Research Scientist - Multi-modal AI & Efficient Generative Models
- Meta
- Redmond, Burlingame, United States
- $180,000 – $250,000
Reality Labs at Meta is building products that make it easier for people to connect with the ones they love most by enabling compelling experiences in novel computing platforms: Smart Glasses and VR headsets. We are a team of experts developing and shipping products advancing the state of the art in both AI and compute. The Core-AI team brings together a team of applied Vision-Language researchers and systems ML experts working on a range of foundational problems in perception, vision-language models, generative models and model optimization. We do a mix of fundamental research and technology development. We are seeking researchers with a passion to bring capabilities of generative vision and language models to resource constrained settings.
Responsibilities
Drive the organization's goal towards relevant machine learning techniques in the area of multi-modal understanding and generation to build & optimize our intelligent systems that improve Meta's products and experiences
* Effectively communicate complex features and systems in detail while advocating for higher product quality and engineering efficiency
* Conduct applied research to advance the state of the art in efficient generative models (Diffusion models, Multi-modal LLMs, Vision Language Action models) and efficient perception models (Scene and Video understanding)
* Apply research to advance Meta's smartglasses and VR product lines
* Advance the state of the art in your problem area by defining and executing research roadmaps over 6-month or longer timeframes
* Collaborate with different cross-functional teams across the globe in research and product
* Present the outcomes of the research findings as papers in top-tier peer-reviewed conferences in the area
Qualifications
PhD in Computer Science or a related field with published projects in the fields of machine learning, Deep learning with a focus on vision-language models
* Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
* 7+ years of experience leading research projects with industry-wide impact
* Proven development skills in Deep Learning, working with PyTorch or TensorFlow
* Experience developing deep learning models or infrastructure in Python or C/C++
* Experience in one or more of the following areas: deep learning, Computer Vision, language models, Machine Learning or artificial intelligence
* First-authored publications at peer-reviewed conferences, e.g. ICLR, ICML, CVPR, ECCV, ICCV, NeurIPS Experience with CPU/GPU and mobile optimization
* Experience solving complex problems and comparing alternative solutions, trade-offs, and diverse points of view to determine a path forward
Skills
- Machine Learning
- Deep Learning
- Generative Models
- Diffusion Models
- Multimodal LLMs
- Python
- PyTorch







