Nuro
Software Engineer, ML Infrastructure Platform
About this role
Nuro is hiring a Software Engineer for its ML Infrastructure Platform team to build and operate the distributed systems that power model training for autonomous driving. You'll design large-scale data pipelines, GPU training orchestration, and agentic ML workflows while ensuring reliability and operational maturity for critical production systems.
What you'll do
- Build and maintain multi-cluster scheduling, orchestration, and GPU training infrastructure across multiple accelerator generations
- Design and operate large-scale data pipelines for batch/streaming ingestion, storage, and high-throughput data generation
- Develop agentic-first ML workflows that are reproducible, introspectable, and easy for autonomy teams to extend
- Instrument critical training and release pipelines with monitoring, alerting, and incident-response practices
- Optimize infrastructure reliability and performance to reduce bottlenecks in autonomy development velocity
- Drive operational maturity through runbooks, observability tools, and failure mode analysis
What they're looking for
- Python programming (required); C++, Go, or similar systems language
- Kubernetes production operations and cluster management
- Distributed systems design, performance analysis, and reliability patterns
- Large-scale data pipeline architecture and implementation
- GPU training internals and multi-node distributed training
- GCP (Google Cloud Platform) infrastructure
- ML workflow orchestration and observability tooling
- On-call and incident response practices
Benefits
- Base salary: $160,360–$240,540 annually
- Annual performance bonus
- Equity compensation
- Competitive benefits package
- Work on cutting-edge autonomous driving technology at scale
- Collaborative environment with strong technical peers
Opens the official application on the employer’s site. No login required.
Nuro
Nuro builds autonomous vehicle platforms and fleet operations systems, with a focus on reliability, safety, and over-the-air update infrastructure. The company is hiring reliability engineers, software engineers, and operations specialists to improve vehicle hardware resilience, enhance system automation, and ensure fleet operational excellence.
View all jobs at NuroLikely interview questions
- Describe a time you debugged a performance bottleneck in a distributed system. What tools and methodology did you use?
- How would you approach designing a data pipeline to support both batch training and real-time reinforcement learning feedback loops?