Clera
Research Engineer, Synthetic Data
About this role
Research Engineer role focused on designing and building synthetic data pipelines to train AI agents, working within a high-performing team of researchers. You'll develop generation methods, validation systems, and quality metrics that advance model capabilities in RL-based AI alignment.
What you'll do
- Build end-to-end synthetic data pipelines transforming domain workflows into structured training tasks for AI agents
- Collaborate with subject-matter experts to develop synthetic tasks across professional and technical domains
- Design task generation methods producing diverse, realistic, and learnable training examples
- Build tooling to mutate, validate, and iteratively improve synthetic task quality
- Analyze model and agent performance on synthetic tasks to identify learning outcomes and failure modes
- Develop metrics quantifying synthetic task diversity, realism, learnability, and overall quality
What they're looking for
- Python programming
- Data pipeline design and implementation
- Synthetic data generation and validation
- Linux and containerization (Docker)
- ML infrastructure and evaluation frameworks
- Reinforcement learning and agentic AI workflows
- Benchmark design for AI agents or LLMs
- Large-scale structured dataset processing
Benefits
- Competitive salary $150,000–$250,000 annually
- Visa sponsorship available
- Remote collaboration across time zones supported
- On-site in San Francisco with Singapore option
- Work at frontier of AI alignment research
- Team of Olympiad medalists and published researchers
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through an end-to-end synthetic data pipeline you've built—what was the generation strategy, and how did you validate quality?
- How do you approach designing synthetic tasks that are both realistic and learnable for AI agents?