HERE Technologies Logo

HERE Technologies

Lead Data Scientist-(VLM-Multimodal AI ,CV, Deep Learning)

Posted 2 Hours Ago
Be an Early Applicant
Hybrid
Mumbai, Maharashtra
Senior level
Hybrid
Mumbai, Maharashtra
Senior level
Lead development of large-scale vision foundation and multimodal models for image and video understanding. Design, train, and optimize ViT-based and VLM architectures, build RAG and retrieval systems using vector databases, manage large datasets and distributed training, and translate research into production-ready scalable pipelines while evaluating model robustness and retrieval performance.
The summary above was generated by AI
What's the role?We are seeking an experienced Lead Data Scientist to drive the development of advanced Vision Foundation Models, Vision-Language Models, and multimodal AI systems for large-scale image and video understanding.You will lead the design, development, and deployment of scalable AI solutions, translating cutting-edge research into real-world applications. Key Responsibilities
  • Design, train, and optimize large-scale vision foundation models across image and video modalities
  • Develop multimodal AI systems using architectures such as Vision Transformers (ViT), SAM, DINOv3, CLIP, and VLMs
  • Apply self-supervised learning, transfer learning, and fine-tuning approaches for downstream tasks
  • Build and enhance Vision-Language Models for visual reasoning and multimodal understanding
  • Develop Retrieval-Augmented Generation (RAG) pipelines and multimodal knowledge retrieval systems
  • Work with embeddings, vector databases, and semantic search frameworks
  • Build scalable pipelines for training, evaluation, and deployment
  • Manage large-scale image, video, and multimodal datasets
  • Optimize distributed training workflows and model performance
  • Translate research into production-ready solutions and explore emerging approaches in multimodal AI and generative AI
  • Evaluate model quality, robustness, and retrieval effectiveness


Who are you?You bring strong expertise in computer vision, foundation models, and multimodal AI systems, along with the ability to deliver scalable solutions from research to production.
  • Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field
  • Extensive experience in deep learning, computer vision, or multimodal AI
  • Strong programming skills in Python and experience with PyTorch
  • Deep understanding of computer vision, Vision Transformers, self-supervised learning, Vision-Language Models, and multimodal systems
  • Hands-on experience with foundation models such as SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models
  • Experience building RAG pipelines, semantic retrieval systems, and working with embeddings and vector databases such as FAISS, Milvus, Pinecone, or Weaviate
  • Experience working with large-scale image and video datasets and distributed training environments
  • Familiarity with GPU acceleration and scalable ML infrastructure
  • Exposure to generative AI, multimodal reasoning systems, or large-scale perception systems
  • Contributions to research, publications, or open-source projects are valued
What Do We Offer?
  • Opportunity to work on cutting-edge AI and multimodal technologies
  • A collaborative, inclusive, and innovation-driven work environment
  • Opportunities to learn, grow, and advance your career
  • Exposure to large-scale, real-world AI challenges and global impact
  • Competitive compensation and performance-based bonus
  • Flexible and hybrid working options
  • Employee wellness programs and professional development support
  HERE Technologies is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, color, religion, gender, gender identity, sexual orientation, age, or disability.
Who are we?

HERE Technologies is a location data and technology platform company. We empower our customers to achieve better outcomes – from helping a city manage its infrastructure or a business optimize its assets to guiding drivers to their destination safely.


At HERE we take it upon ourselves to be the change we wish to see. We create solutions that fuel innovation, provide opportunity and foster inclusion to improve people’s lives. If you are inspired by an open world and driven to create positive change, join us. Learn more about us on our YouTube Channel.




About the TeamYou will be part of a highly collaborative AI/ML team focused on developing next-generation Vision Foundation Models (VFMs), Vision-Language Models (VLMs), and multimodal AI systems. The team works at the intersection of research and scalable production systems, driving innovation in large-scale image and video understanding.

Similar Jobs at HERE Technologies

2 Hours Ago
Hybrid
Senior level
Senior level
Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Integrate, validate, and optimize real-time edge road perception systems on automotive and embedded hardware. Develop validation, benchmarking, and regression testing pipelines for CV tasks, profile performance (latency, FPS, memory, GPU/CPU), deploy models with TensorRT/ONNX/CUDA/OpenCV, and troubleshoot integration and hardware issues.
Top Skills: ArmC++CudaLinuxNvidia JetsonOnnx RuntimeOpencvPythonQualcomm SnapdragonTensorrt
Senior level
Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Lead and scale a global Quality Assurance & Testing organization to define quality strategy, validation methodologies, automation-first practices, and KPI-based quality measurement. Build data-driven validation pipelines, predictive quality models, and independent end-to-end testing to ensure release readiness. Partner with Product, Engineering, and Operations to influence prioritization and deliver measurable improvements in product quality.
Top Skills: Ai/MlAnomaly DetectionAutomation FrameworksData PipelinesETLPredictive Quality ModelsPythonSQL
Yesterday
Hybrid
Senior level
Senior level
Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Lead design, development, and operation of scalable microservices on AWS using Java/Scala/Python. Own end-to-end feature delivery, produce technical designs, mentor peers, and drive engineering best practices. Collaborate with global stakeholders and use data-driven decisions to recommend architectural changes.
Top Skills: AWSJavaMicroservicesPythonScala

What you need to know about the Kolkata Tech Scene

When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account