JobConnect

Research Scientist, Multi-Modal Human Understanding

  • Meta
  • Pittsburgh, Burlingame, United States
  • $180,000 – $250,000

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Responsibilities
Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
* Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
* Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
* Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
* Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
* Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation

Qualifications
Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
* 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
* 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
* Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
* Experience writing production-quality or research-quality code in Python for multi-modal AI applications Experience developing or fine-tuning Vision-Language Models for human understanding tasks
* Experience with video foundation models, temporal transformers, or large-scale video pretraining
* Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
* Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation

Skills

  • Machine Learning
  • Deep Learning
  • Vision-Language Models
  • PyTorch
  • Computer Vision
  • Natural Language Processing
  • Research

Related jobs

MetaApply for this job