Skip to main content

Scale AI

Machine Learning Research Scientist, Evaluations

San Francisco, CA; Seattle, WA; New York, NYFrom $225.8kfull-timemidAdded today

About this role

Scale is hiring a Machine Learning Research Scientist to develop evaluation frameworks and diagnostic methods for frontier LLMs and multimodal models. You'll identify failure modes, design benchmarks, and translate findings into actionable insights for leading AI labs.

What you'll do

  • Analyze and diagnose failure modes in frontier LLMs and agents across reasoning, robustness, and alignment
  • Design benchmarks and evaluation methods for text and multimodal model capabilities
  • Apply post-training expertise (SFT, RLHF, reward modeling) to connect failures to training interventions
  • Publish research findings at top-tier AI conferences
  • Collaborate with foundation model labs to inform next-generation AI development
  • Perform root cause analysis on model behavior and capability gaps

What they're looking for

  • Deep learning and reinforcement learning
  • LLM post-training techniques (RLHF, preference modeling, instruction tuning)
  • LLM evaluation and benchmark design
  • Large-scale model fine-tuning
  • Multimodal AI systems
  • Research communication and publication
  • Python and ML frameworks
  • Model failure analysis and diagnostics

Benefits

  • Comprehensive health, dental, and vision coverage
  • Equity-based compensation
  • Retirement benefits
  • Learning and development stipend
  • Generous paid time off
  • Commuter stipend eligibility
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Scale AI

Scale AI builds a Generative AI Data Engine and ML infrastructure platforms that power LLM training, evaluation, and production serving at scale, along with data solutions for robotics and autonomous driving. The company is hiring Software Engineers for full-stack feature development and infrastructure systems, identity/security specialists for platform engineering, and Solutions Engineers to support enterprise clients and pre-sales processes.

Website
scale.com
View all jobs at Scale AI

Likely interview questions

  • Describe a time you diagnosed a failure mode in a large language model—what was your approach and what did you learn?
  • How would you design an evaluation benchmark to measure reasoning capabilities in both text and multimodal models?