Clera
Research Engineer, Synthetic Data
About this role
Join an early-stage AI infrastructure startup to design and build end-to-end synthetic data pipelines for reinforcement learning environments. You'll transform domain expertise into realistic, scalable training tasks and develop metrics to evaluate their effectiveness for AI agent learning.
What you'll do
- Build and maintain synthetic data pipelines that convert domain workflows into structured training tasks
- Collaborate with subject-matter experts to design high-quality synthetic tasks across professional domains
- Engineer task generation methods that produce diverse, realistic, and learnable data at scale
- Create tooling for mutation, validation, and iterative improvement of synthetic tasks
- Analyze agent performance on synthetic tasks to identify what they teach and failure modes
- Develop metrics quantifying task diversity, realism, learnability, and overall quality
What they're looking for
- Python proficiency
- Docker and Linux environments
- Synthetic data research methods
- End-to-end pipeline design and development
- Environment, evaluation, and benchmark design
- First-principles reasoning about task design and scoring functions
- Data quality assurance and edge case identification
- Async collaboration and technical communication
Benefits
- Competitive salary: $150,000–$250,000 annually
- Equity participation in well-funded early-stage AI company
- Visa sponsorship available
- Work with Olympiad medalists and published researchers
- On-site office environment in San Francisco
- Opportunity to shape foundational AI infrastructure
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a synthetic data pipeline you've built end-to-end—what made the data 'good' and how did you validate it?
- How do you approach designing synthetic tasks that are both realistic and learnable for AI agents?