Skip to main content

Clera

Research Engineer, Synthetic Data

San Francisco$150k–$250kfulltimemidAdded today

About this role

Join a 15-person research engineering team building synthetic data infrastructure for AI agent training. You'll own the end-to-end pipeline that converts real-world workflows into high-quality, diverse training tasks using reinforcement learning and post-training methodologies.

What you'll do

  • Build and maintain the synthetic data pipeline transforming domain workflows into training tasks for AI agents
  • Collaborate with subject-matter experts to generate synthetic tasks across professional and technical domains
  • Design task generation methods producing diverse, realistic, and learnable outputs at scale
  • Create tooling to validate, mutate, and iteratively improve synthetic tasks
  • Analyze agent performance on synthetic tasks to identify learning gaps and failure modes
  • Develop metrics quantifying task diversity, realism, learnability, and quality

What they're looking for

  • Python programming
  • Data pipeline design and implementation
  • Synthetic data generation and evaluation
  • Docker and Linux containerization
  • ML/AI infrastructure development
  • Evaluation frameworks and benchmarking
  • Reinforcement learning paradigms
  • Large-scale dataset processing and validation

Benefits

  • Visa sponsorship available
  • Equity participation in early-stage startup
  • High ownership with minimal bureaucracy
  • Work with published researchers and Olympiad medalists
  • Unstructured problem-solving environment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Walk us through a synthetic data pipeline you've built—what made it challenging and how did you measure its quality?
  • Describe a time you discovered data quality issues in a generated or algorithmically produced dataset. How did you detect and fix them?