Clera
Research Engineer, Synthetic Data
About this role
Build synthetic data pipelines that transform domain workflows into scalable AI training tasks. Work with a selective team of researchers and engineers to design generation methods, validation systems, and quality metrics that enable AI agents to learn complex behaviors.
What you'll do
- Design and implement end-to-end synthetic data pipelines converting domain workflows into structured training tasks
- Collaborate with subject-matter experts to generate realistic synthetic tasks across professional and technical domains
- Develop scalable task generation methods producing diverse, realistic, and learnable outputs
- Build automated tooling for task mutation, validation, and continuous quality improvement
- Analyze model and agent performance on synthetic tasks to identify learning patterns and failure modes
- Define and implement metrics for task diversity, realism, learnability, and overall data quality
What they're looking for
- Python and Linux development
- Docker and containerization
- ML infrastructure and data pipeline design
- Synthetic data systems and methodology
- Evaluation frameworks and benchmarking for AI agents
- Reinforcement learning and agentic AI workflows
- LLM post-training pipelines
- First-principles reasoning and edge-case detection
Benefits
- Visa sponsorship available
- Competitive salary range $150,000–$250,000
- Work with Olympiad medalists and published researchers
- Direct ownership over generation methods and validation systems
- On-site collaboration in San Francisco (Singapore candidates considered)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a synthetic data pipeline you've built end-to-end—what made it effective for training models?
- How do you approach measuring synthetic data quality, and what trade-offs exist between diversity, realism, and learnability?