Mercor
Research Engineer - Environments, Data and Post-Training
About this role
Mercor is seeking a Research Engineer to enhance language model performance through post-training and data generation methods. This role involves collaborating with researchers to design experiments and build systems that optimize LLM capabilities in real-world contexts.
What you'll do
- Develop post-training pipelines and evaluate their effectiveness
- Conduct reward-shaping experiments and improve algorithmic methods
- Assess data usability and its impact on performance benchmarks
- Create scalable data generation and augmentation systems
- Design rubrics and scoring frameworks for evaluation
- Collaborate with AI teams and experts on training data
What they're looking for
- Applied research with a focus on model evaluation
- Strong coding skills in machine learning
- Understanding of data structures and algorithms
- Familiarity with APIs and databases
- Ability to analyze model behavior and data quality
- Experience in a fast-paced research environment
- Knowledge of synthetic data generation
- Experience with LLM evaluations
Benefits
- Bi-annual performance bonuses
- Generous equity grant
- Relocation bonus up to $15k
- $10K housing bonus near the office
- $1.5K monthly meal stipend
- Free Equinox membership
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mercor
Mercor builds a marketplace platform connecting expert talent to AI opportunities, supported by identity infrastructure, matching algorithms, and internal tools for data management. The company is hiring Software Engineers, Machine Learning Engineers, Fullstack Engineers, and Security Engineers to develop backend systems, ML models, cloud infrastructure, and distributed platforms.
- Website
- mercor.io
Likely interview questions
- Walk us through a post-training or RLHF project you've worked on. What were the key challenges in reward design, and how did you measure whether your approach actually improved model performance?
- How would you approach designing and validating a synthetic data generation pipeline for LLM training? What metrics would you use to quantify data quality and usability?