Skip to main content

Clera

Research Engineer, Benchmarks

San Francisco$150k–$250kfulltimemidAdded today

About this role

Design and build benchmarks that rigorously evaluate frontier AI agents on realistic, domain-specific tasks. You'll own benchmark quality, work with domain experts to define evaluation criteria, operate evaluation infrastructure at scale, and validate that benchmarks meaningfully predict real-world performance.

What you'll do

  • Design, implement, and maintain internal benchmarks for evaluating AI agents on domain-specific workflows
  • Collaborate with subject-matter experts to translate realistic workflows into well-scoped evaluation tasks and success criteria
  • Build and operate scalable infrastructure to run models and agents against benchmark tasks
  • Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes
  • Validate benchmark performance correlates with real-world outcomes and frontier lab expectations
  • Write technical documentation and benchmark reports for research and engineering audiences

What they're looking for

  • Python programming
  • Docker and Linux systems administration
  • AI benchmark design and evaluation methodology
  • Statistical analysis and validation
  • Agent and LLM evaluation experience
  • Metric design and quality assurance
  • Technical writing and communication
  • Infrastructure and systems design

Benefits

  • Competitive salary range: $150,000–$250,000 USD annually
  • Equity compensation
  • Visa sponsorship available
  • On-site in San Francisco with in-person collaboration
  • High-ownership role on a small, technical team
  • Work on frontier AI evaluation problems
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Walk us through a benchmark or evaluation system you've designed—how did you ensure it was both realistic and reliable?
  • Describe your experience developing metrics to assess benchmark quality. How did you validate that your metrics actually predicted real-world performance?