Mercor
Research Engineer - Environments, Data and Post-Training
About this role
Mercor seeks a Research Engineer to develop and implement post-training methods for frontier AI models. You'll design experiments across datasets, reward functions, and optimization strategies while building scalable pipelines that measure data quality and drive model improvements.
What you'll do
- Implement novel post-training methods to improve model reasoning, tool use, and agentic behavior
- Design and execute experiments on training recipes, reward functions, and optimization strategies like GRPO and DAPO
- Build reinforcement learning pipelines at scale, including RLVR systems
- Create methods for measuring data quality, usability, and causal impact on model performance
- Develop rubrics, evaluators, and benchmarks that inform training decisions
- Investigate model capabilities and failure modes to develop targeted training interventions
What they're looking for
- Machine learning model training and evaluation
- Reinforcement learning and post-training methods
- Python and ML system implementation
- Experimental design and rigorous analysis
- Data-centric ML and synthetic data generation
- Distributed systems and cloud infrastructure
- Knowledge of current AI research landscape
- Large-scale evaluation and data pipeline infrastructure
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15,000 relocation bonus
- $10,000 housing bonus for proximity to office
- Health, dental, and vision insurance
- Free Equinox membership, meal stipend, wellness reimbursement
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mercor
Mercor builds a marketplace platform connecting expert talent to AI opportunities, supported by identity infrastructure, matching algorithms, and internal tools for data management. The company is hiring Software Engineers, Machine Learning Engineers, Fullstack Engineers, and Security Engineers to develop backend systems, ML models, cloud infrastructure, and distributed platforms.
- Website
- mercor.io
Likely interview questions
- Describe your experience with post-training methods like GRPO or DAPO and how you've applied them to improve model behavior.
- How have you designed experiments to measure the causal impact of data or training interventions on model performance?