Clera
Research Engineer, Benchmarks
About this role
Design and build rigorous benchmarks for evaluating frontier AI agents on realistic, domain-specific workflows. You'll own evaluation quality and infrastructure at a startup focused on AI assessment, working closely with subject-matter experts to create benchmarks that frontier labs and enterprises trust.
What you'll do
- Design, implement, and maintain internal benchmarks for evaluating AI agents on domain-specific tasks
- Partner with subject-matter experts to translate real-world workflows into well-scoped evaluation tasks
- Build and operate scalable infrastructure to run models and agents against benchmark tasks
- Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes
- Validate that benchmark performance correlates with real-world agent behavior and customer needs
- Write documentation and reports that communicate results clearly to technical audiences
What they're looking for
- Python programming
- Docker and containerization
- Linux environments
- AI agent and LLM evaluation design
- Benchmark infrastructure and scale operations
- Statistical analysis and metrics development
- Technical writing and cross-functional communication
- Problem-solving in unstructured environments
Benefits
- Visa sponsorship available
- Work on cutting-edge AI evaluation at a startup
- Competitive salary range of $150,000–$250,000 annually
- On-site in San Francisco
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a benchmark or evaluation project you designed—what made it rigorous and how did you validate its real-world relevance?
- Describe your experience building infrastructure to run AI models or agents at scale. What challenges did you encounter and how did you solve them?