Cartesia
Software Engineer, Data Infrastructure
About this role
Cartesia seeks a Software Engineer to lead data infrastructure strategy and execution, building scalable pipelines for multimodal datasets (text, audio, video) that power foundation models. You'll design systems for data acquisition, processing, and curation while partnering with research and inference teams to ensure data quality directly impacts model capabilities.
What you'll do
- Define multi-modal data strategy across pre-training and post-training with focus on audio, synthetic, and web-scale sources
- Design and operate scalable, high-throughput data pipelines covering ingestion, preprocessing, augmentation, versioning, and GPU-aware training data loading
- Partner with research and inference teams to co-design data systems with training and serving infrastructure
- Establish data quality standards with feedback loops between dataset characteristics and model behavior
- Source novel datasets and manage relationships and budgets with external data vendors and partners
What they're looking for
- ML data infrastructure and training pipelines
- Multimodal data handling (audio formats, preprocessing, augmentation)
- Large-scale data loading and dataset versioning
- Python or similar languages with strong engineering practices
- Generative model dataset building and evaluation
- Cross-functional technical leadership in research environments
- Data systems architecture for model training and inference
- Cloud data storage and streaming patterns
Benefits
- Competitive base salary with equity package
- Fully covered medical, dental, and vision insurance for family
- 9 weeks paternity and 12 weeks maternity leave
- 401(k) retirement plan
- Monthly commuter allowance
- Flexible PTO and provided meals and snacks
Opens the official application on the employer’s site. No login required.
Cartesia
Cartesia builds multimodal AI models and voice AI platforms that power enterprise applications. The company is hiring for roles spanning customer support, enterprise deployments, inference infrastructure, internal developer tooling, and forward-deployed engineering to scale their AI solutions across production environments.
- Website
- cartesia.ai
Likely interview questions
- Walk us through your experience building data pipelines for machine learning models—what scale and complexity did you work with?
- How have you approached data quality measurement and its relationship to model performance in past roles?