Research Scientist - Multi-modal AI & Efficient Generative Models
Description
Reality Labs at Meta is building products that make it easier for people to connect with the ones they love most by enabling compelling experiences in novel computing platforms: Smart Glasses and VR headsets. We are a team of experts developing and shipping products advancing the state of the art in both AI and compute. The Core-AI team brings together a team of applied Vision-Language researchers and systems ML experts working on a range of foundational problems in perception, vision-language models, generative models and model optimization. We do a mix of fundamental research and technology development. We are seeking researchers with a passion to bring capabilities of generative vision and language models to resource constrained settings.
Responsibilities
Drive the organization's goal towards relevant machine learning techniques in the area of multi-modal understanding and generation to build & optimize our intelligent systems that improve Meta's products and experiences Effectively communicate complex features and systems in detail while advocating for higher product quality and engineering efficiency Conduct applied research to advance the state of the art in efficient generative models (Diffusion models, Multi-modal LLMs, Vision Language Action models) and efficient perception models (Scene and Video understanding) Apply research to advance Meta's smartglasses and VR product lines Advance the state of the art in your problem area by defining and executing research roadmaps over 6-month or longer timeframes Collaborate with different cross-functional teams across the globe in research and product Present the outcomes of the research findings as papers in top-tier peer-reviewed conferences in the area
Qualifications
PhD in Computer Science or a related field with published projects in the fields of machine learning, Deep learning with a focus on vision-language models Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 7+ years of experience leading research projects with industry-wide impact Proven development skills in Deep Learning, working with PyTorch or TensorFlow Experience developing deep learning models or infrastructure in Python or C/C++ Experience in one or more of the following areas: deep learning, Computer Vision, language models, Machine Learning or artificial intelligence First-authored publications at peer-reviewed conferences, e.g. ICLR, ICML, CVPR, ECCV, ICCV, NeurIPS Experience with CPU/GPU and mobile optimization Experience solving complex problems and comparing alternative solutions, trade-offs, and diverse points of view to determine a path forward
Compensation: $184,000/year to $257,000/year + bonus + equity + benefits