Clera
Research Engineer, Benchmarks
About this role
Design and build rigorous benchmarks to evaluate frontier AI agents on realistic domain-specific tasks. You'll own the quality of evaluations trusted by leading labs and enterprise customers, ensuring they accurately measure agent performance in real-world workflows.
What you'll do
- Design and implement internal benchmarks for evaluating frontier agents on domain-specific tasks
- Partner with subject-matter experts to define realistic workflows and evaluation criteria
- Build scalable infrastructure to run models and agents against benchmark tasks
- Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes
- Validate that benchmark performance correlates with real-world evaluations and customer needs
- Write clear documentation and reports to communicate results credibly to technical audiences
What they're looking for
- Python programming
- Docker and containerization
- Linux environments
- AI/ML evaluation and benchmarking
- Benchmark and environment design
- Data analysis and metrics development
- Technical writing and documentation
- First-principles reasoning about task design and scoring
Benefits
- Salary: $150,000–$250,000 annually
- Visa sponsorship available
- Early-stage equity opportunity
- Significant career growth potential
- Work with elite technical team including Olympiad medalists and published researchers
- On-site position in San Francisco with collaborative, research-driven culture
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a benchmark or evaluation you've designed—how did you ensure it was realistic and reliable?
- Describe your experience building infrastructure to run models or agents at scale. What challenges did you face?