Skip to main content

Cartesia

Research Engineer, Data Infrastructure (Language Modeling)

*HQ - San Francisco, CA$200k–$350kfulltimemidAdded today

About this role

Cartesia seeks a Research Engineer to build scalable data infrastructure powering their language model pretraining. You'll design and operate high-throughput pipelines for dataset acquisition, processing, and curation, directly optimizing model quality through data-driven experimentation and close collaboration with research teams.

What you'll do

  • Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets
  • Design and run scalable, reproducible data pipelines covering ingestion, preprocessing, filtering, deduplication, and augmentation
  • Conduct ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality
  • Partner with research and infrastructure teams to co-design data loading, versioning, and experimentation systems
  • Establish and enforce data quality standards with tight feedback loops between dataset characteristics and model performance
  • Identify and source novel datasets; manage relationships and budgets with external data vendors

What they're looking for

  • ML data infrastructure and training data pipelines
  • Dataset versioning and large-scale data loading
  • Modern software engineering practices and testing
  • Generative model evaluation and understanding
  • Large-scale data processing (Ray, Spark, or Kubernetes preferred)
  • Language model pretraining and inference knowledge
  • Data quality assessment and standards enforcement
  • Pipeline design and reproducibility

Benefits

  • Competitive base salary with equity package
  • Fully covered medical, dental, and vision insurance for you and family
  • Flexible PTO
  • 9 weeks paternity and 12 weeks maternity leave
  • 401(k) and monthly commuter allowance
  • Daily meals, snacks, and collaborative office culture
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Cartesia

Cartesia builds multimodal AI models and voice AI platforms that power enterprise applications. The company is hiring for roles spanning customer support, enterprise deployments, inference infrastructure, internal developer tooling, and forward-deployed engineering to scale their AI solutions across production environments.

View all jobs at Cartesia

Likely interview questions

  • Walk us through a data pipeline you've built for ML training—what were the bottlenecks and how did you optimize them?
  • How do you approach data versioning and reproducibility in large-scale training pipelines?