Skip to main content

Clera

Research Engineer, Benchmarks

San Francisco$150k–$250kfulltimemidAdded today

About this role

Design and build rigorous benchmarks for evaluating frontier AI agents on realistic, domain-specific workflows. You'll own evaluation quality and infrastructure at a startup focused on AI assessment, working closely with subject-matter experts to create benchmarks that frontier labs and enterprises trust.

What you'll do

  • Design, implement, and maintain internal benchmarks for evaluating AI agents on domain-specific tasks
  • Partner with subject-matter experts to translate real-world workflows into well-scoped evaluation tasks
  • Build and operate scalable infrastructure to run models and agents against benchmark tasks
  • Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes
  • Validate that benchmark performance correlates with real-world agent behavior and customer needs
  • Write documentation and reports that communicate results clearly to technical audiences

What they're looking for

  • Python programming
  • Docker and containerization
  • Linux environments
  • AI agent and LLM evaluation design
  • Benchmark infrastructure and scale operations
  • Statistical analysis and metrics development
  • Technical writing and cross-functional communication
  • Problem-solving in unstructured environments

Benefits

  • Visa sponsorship available
  • Work on cutting-edge AI evaluation at a startup
  • Competitive salary range of $150,000–$250,000 annually
  • On-site in San Francisco
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Walk us through a benchmark or evaluation project you designed—what made it rigorous and how did you validate its real-world relevance?
  • Describe your experience building infrastructure to run AI models or agents at scale. What challenges did you encounter and how did you solve them?