Mercor
Research Engineer – Benchmarking, Evals & Failure Analysis
About this role
Mercor is seeking a Research Engineer to enhance AI models through benchmarking, evaluation systems, and failure analysis. This role involves collaborating with teams to define metrics and improve data quality in a fast-paced environment in San Francisco.
What you'll do
- Design and maintain benchmarking metrics for various AI behaviors
- Develop and manage evaluation systems for tracking model performance
- Conduct failure analysis on model outputs to identify improvement areas
- Create and refine rubrics and scoring frameworks for evaluations
- Assess data quality and impact on benchmarks to guide data strategies
- Collaborate with teams to align evaluations with training goals
What they're looking for
- Background in applied research and model evaluation
- Strong coding skills related to ML models
- Familiar with data structures and algorithms
- Experience with APIs and SQL/NoSQL databases
- Ability to analyze model behavior and evaluate data quality
- Willingness to work in-office in a dynamic setting
Benefits
- Bi-annual performance bonuses
- Equity grant vested over 4 years
- Relocation bonuses up to $15k
- Housing bonuses for nearby residents
- $1.5k monthly meal stipend
- Free Equinox membership
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mercor
Mercor builds a marketplace platform connecting expert talent to AI opportunities, supported by identity infrastructure, matching algorithms, and internal tools for data management. The company is hiring Software Engineers, Machine Learning Engineers, Fullstack Engineers, and Security Engineers to develop backend systems, ML models, cloud infrastructure, and distributed platforms.
- Website
- mercor.io
Likely interview questions
- Walk us through a benchmarking or evaluation system you've built for LLMs. How did you design the metrics, and how did you validate that they were measuring what mattered?
- Tell us about a time you identified a systematic failure mode in model outputs. How did you categorize it, quantify its prevalence, and use those insights to drive improvements?