Skip to main content

DeepMind

Research Engineer, Human Understanding

Los Angeles, California, US; Mountain View, California, USmidAdded 1 month ago

About this role

Google DeepMind is hiring a Research Engineer (L5) to develop multimodal AI models for human understanding, focusing on speech, audio, and visual data. You'll design foundational capabilities to understand and generate human representations while building defenses against deepfakes and impersonation within the Frontier AI unit.

What you'll do

  • Research and implement novel multimodal models for holistic human understanding across visual, audio, and textual data
  • Conduct experimental research cycles from hypothesis through deployment
  • Lead substantial technical projects from ideation to evaluation with cross-functional collaboration
  • Develop scalable research infrastructure for multimodal models and datasets
  • Design and execute strategies for tuning vision language models and foundation models for specific tasks
  • Drive technical direction for core components addressing complex, ambiguous problems

What they're looking for

  • Machine learning model development (audio, speech, visual models)
  • Vision language models and foundation model tuning
  • Python programming and deep learning frameworks (e.g., JAX)
  • Independent research design, implementation, and analysis
  • Multimodal learning integrating vision, audio, and text
  • Generative AI techniques and architectures
  • Reinforcement learning or alignment methods
  • Privacy-preserving machine learning and responsible AI

Benefits

  • Competitive salary range: $174,000 - $252,000 USD
  • Bonus and equity compensation
  • Health and wellness benefits package
  • Opportunity to contribute to AGI research at Google DeepMind
  • Work on impactful, responsible AI deployment across Google products
  • Collaboration with leading AI researchers and scientists
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

DeepMind

DeepMind develops advanced AI systems across multiple domains, including multimodal models for understanding human communication and detecting manipulated media, as well as AI applications in materials science discovery. The company is hiring research engineers and researchers to build foundational AI capabilities, create defenses against digital misinformation, and accelerate scientific discovery through machine learning and computational innovation.

View all jobs at DeepMind

Likely interview questions

  • Walk us through a multimodal machine learning project you've led—specifically one involving audio, speech, or visual data. What were the key technical challenges and how did you address them?
  • Describe your experience tuning and adapting vision language models (VLMs) for specific tasks. What techniques have you found most effective, and how did you evaluate their performance?