Lead Data Scientist-(VLM-Multimodal AI ,CV, Deep Learning)

HERE TechnologiesApplyPublished 8 days agoFirst seen 1 days ago
Apply

Responsibilities

We are seeking an experienced Lead Data Scientist to drive the development of advanced Vision Foundation Models, Vision-Language Models, and multimodal AI systems for large-scale image and video understanding.

You will lead the design, development, and deployment of scalable AI solutions, translating cutting-edge research into real-world applications.

Key Responsibilities

  • Design, train, and optimize large-scale vision foundation models across image and video modalities
  • Develop multimodal AI systems using architectures such as Vision Transformers (ViT), SAM, DINOv3, CLIP, and VLMs
  • Apply self-supervised learning, transfer learning, and fine-tuning approaches for downstream tasks
  • Build and enhance Vision-Language Models for visual reasoning and multimodal understanding
  • Develop Retrieval-Augmented Generation (RAG) pipelines and multimodal knowledge retrieval systems
  • Work with embeddings, vector databases, and semantic search frameworks
  • Build scalable pipelines for training, evaluation, and deployment
  • Manage large-scale image, video, and multimodal datasets
  • Optimize distributed training workflows and model performance
  • Translate research into production-ready solutions and explore emerging approaches in multimodal AI and generative AI
  • Evaluate model quality, robustness, and retrieval effectiveness

Qualifications

You bring strong expertise in computer vision, foundation models, and multimodal AI systems, along with the ability to deliver scalable solutions from research to production.

  • Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field
  • Extensive experience in deep learning, computer vision, or multimodal AI
  • Strong programming skills in Python and experience with PyTorch
  • Deep understanding of computer vision, Vision Transformers, self-supervised learning, Vision-Language Models, and multimodal systems
  • Hands-on experience with foundation models such as SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models
  • Experience building RAG pipelines, semantic retrieval systems, and working with embeddings and vector databases such as FAISS, Milvus, Pinecone, or Weaviate
  • Experience working with large-scale image and video datasets and distributed training environments
  • Familiarity with GPU acceleration and scalable ML infrastructure
  • Exposure to generative AI, multimodal reasoning systems, or large-scale perception systems
  • Contributions to research, publications, or open-source projects are valued

What Do We Offer?

  • Opportunity to work on cutting-edge AI and multimodal technologies
  • A collaborative, inclusive, and innovation-driven work environment
  • Opportunities to learn, grow, and advance your career
  • Exposure to large-scale, real-world AI challenges and global impact
  • Competitive compensation and performance-based bonus
  • Flexible and hybrid working options
  • Employee wellness programs and professional development support

HERE Technologies is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, color, religion, gender, gender identity, sexual orientation, age, or disability.