Skip to main content

Mercor

Research Engineer – Benchmarking

San Francisco$130k–$500kfulltimemidAdded today

About this role

Mercor seeks a Research Engineer to design and operate benchmarking pipelines and evaluation systems for frontier language models. You'll own end-to-end eval infrastructure, conduct failure analysis, and collaborate with AI research teams to measure tool use, agentic behavior, and real-world reasoning.

What you'll do

  • Design and maintain benchmarks for tool use, agentic behavior, and real-world reasoning that scale with training
  • Build and operate LLM evaluation systems including runs, scoring, dashboards, and performance tracking at scale
  • Conduct systematic failure analysis on model outputs and categorize failure modes to inform training improvements
  • Create rubrics, automated evaluators, and scoring frameworks balancing rigor with scalability
  • Quantify data quality and usability impact on key benchmarks to guide data curation
  • Collaborate across AI research, applied teams, and data producers to align evaluations with training objectives

What they're looking for

  • Applied research in model evaluation, benchmarking, or failure analysis
  • Strong coding and hands-on experience with ML models and evaluation frameworks
  • Data structures, algorithms, and backend systems design
  • APIs, SQL/NoSQL databases, and cloud platforms
  • Reasoning about model behavior and experimental analysis
  • LLM evaluation pipeline development
  • Synthetic data generation or rubric design
  • RL-style workflows using evals for reward shaping

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15,000 relocation bonus
  • $10,000 housing bonus (if within 0.5 miles of office)
  • $1,500 monthly meals stipend
  • Free Equinox membership, health/dental/vision insurance
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Mercor

Mercor builds a marketplace platform connecting expert talent to AI opportunities, supported by identity infrastructure, matching algorithms, and internal tools for data management. The company is hiring Software Engineers, Machine Learning Engineers, Fullstack Engineers, and Security Engineers to develop backend systems, ML models, cloud infrastructure, and distributed platforms.

Website
mercor.io
View all jobs at Mercor

Likely interview questions

  • Walk us through a benchmarking project you've led—how did you design metrics and ensure they stayed aligned with research goals as the model evolved?
  • Describe your experience building or operating LLM evaluation systems. What were the biggest scalability or quality challenges you faced?