Clera
Research Engineer
About this role
Join an AI infrastructure company building systems for training and evaluating frontier AI agents. As a Research Engineer, you'll design benchmarks, synthetic data pipelines, and quality control automation while working alongside world-class researchers in a high-ownership, early-stage environment.
What you'll do
- Build systems for creating training environments, improving data quality, and converting real-world workflows into tasks and benchmarks
- Design experiments to understand model behavior, agent failure modes, and identify data quality issues
- Develop tools for researchers and data vendors to create higher-quality tasks, trajectories, and feedback loops
- Manage the full lifecycle of agent training data from task design through trajectory collection, evaluation, and validation
- Partner with external vendors to identify bottlenecks and improve data engine quality and throughput
- Build metrics and analyses to assess task and environment usefulness for training frontier agents
What they're looking for
- Python
- Docker
- Linux
- Benchmark design and evaluation
- Data pipeline architecture
- Metrics and validation methodology
- Reinforcement learning (preferred)
- Independent infrastructure building
Benefits
- Competitive salary ($150K–$250K annually)
- Equity participation in well-resourced AI company
- Visa sponsorship available
- Work with International Olympiad medalists and top-tier researchers
- High ownership and direct impact on frontier AI development
- On-site collaborative environment in San Francisco
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a time you built research infrastructure or tooling with minimal guidance—what was the hardest part and how did you solve it?
- Describe your experience designing benchmarks or evaluation frameworks. How did you ensure they actually measured what mattered for model training?