Wayve
Software Engineer, Data Flywheel Platform
SunnyvalemidAdded today
About this role
Wayve seeks a senior software engineer to build the platform infrastructure powering its data flywheel and foundation-model stack for autonomous driving. You'll design and scale the pipelines, systems, and self-serve products that enable teams to process world-scale fleet data, evaluate foundation models, and train large-scale AI systems.
What you'll do
- Build data curation and enrichment pipelines that transform fleet data into high-signal training data at scale
- Develop evaluation infrastructure for foundation-model progress, including offline and closed-loop evaluation harnesses
- Design and optimize distributed training and serving infrastructure for large pretrained models
- Construct the data-platform backbone using distributed processing, vector search, lakehouse formats, and dataset versioning
- Convert ad-hoc processes into self-serve, observable products with strong testing and monitoring
- Partner with Applied Scientists and ML Engineers to move research prototypes to production deployment
What they're looking for
- Production Python (services, APIs, large-scale data processing)
- Distributed systems and batch/streaming pipelines
- Workflow orchestration tools (Flyte, Airflow, Dagster)
- Data platform technologies (Spark, Ray Data, Daft, Databricks)
- Vector search and embedding systems
- Large codebase ownership and software architecture
- Observability and testing practices
- Model training and serving infrastructure
Opens the official application on the employer’s site. No login required.
Wayve
Wayve develops autonomous driving AI technology and systems for vehicles. The company is hiring for validation engineers, systems engineers, ML engineers, integration specialists, and program managers to build and deploy autonomous driving platforms across testing, AI development, hardware integration, and production scaling.
- Website
- wayve.ai
Likely interview questions
- Describe your experience building and scaling production data pipelines at large scale—what were the key challenges and how did you address them?
- How have you approached converting one-off scripts or manual processes into reliable, self-serve systems that other teams can operate?