Waymo
2027 Summer Intern, MS/PhD, Software/ML Engineer, Simulation
Mountain View, California, USA$145.6k–$176.8kinternshipinternAdded today
About this role
Waymo seeks MS/PhD interns to develop multimodal data pipelines and Vision-Language Model workflows for autonomous driving simulation and testing. You'll design data systems combining sensor inputs and 3D maps, build ground-truth datasets, and integrate automated labeling into simulation pipelines.
What you'll do
- Design and build multimodal data pipelines in Python or C++ integrating camera, LiDAR, and 3D map data
- Prototype and iterate on Vision-Language Model workflows to detect and categorize driving scenarios
- Create ground-truth evaluation datasets and measure model precision/recall
- Integrate automated labeling pipelines into data storage and simulation systems
- Improve detection accuracy through systematic evaluation and iteration
What they're looking for
- Python or C++
- Data pipeline design
- Computer vision
- Vision-Language Models (VLMs)
- SQL and data analysis
- Statistics and evaluation metrics
- 3D geometry and sensor processing
- Prompt engineering
Benefits
- Competitive hourly compensation ($70/hr for Masters, $85/hr for PhD)
- Housing/relocation bonus (if applicable)
- Medical, dental, and vision insurance
- Free meals (breakfast, lunch, dinner, snacks) onsite
- Free Google shuttle access
- Onsite gym and intern networking events
Opens the official application on the employer’s site. No login required.
Waymo
Waymo develops autonomous driving technology and vehicles, building the AI systems, simulation platforms, and infrastructure that power the Waymo Driver. The company is hiring for ML infrastructure engineers, platform engineers, labeling system developers, backend software engineers, and automotive systems engineers to scale its autonomous driving capabilities.
- Website
- waymo.com
Likely interview questions
- Describe your experience building or working with multimodal data pipelines. What challenges did you encounter?
- Tell us about a time you worked with Vision-Language Models or prompt engineering—how did you approach iterating on prompts?