Zoox
Software Engineer - Planner GPU Compute
About this role
Zoox seeks a GPU performance expert to optimize motion planning algorithms for autonomous vehicles. You'll analyze and enhance GPU-based compute across ML inference and custom CUDA kernels, ensuring low-latency, resource-efficient operation critical to safe self-driving functionality.
What you'll do
- Identify and eliminate GPU performance bottlenecks through metrics analysis and profiling
- Adapt motion planner algorithms to support multiple GPU architectures with varying resources
- Optimize machine learning models for latency and GPU memory constraints
- Debug and optimize CUDA kernels using specialized tools like Nvidia Nsight
- Serve as GPU/CUDA technical expert supporting broader Planner team
- Collaborate on multiprocess system performance in robotic/autonomous contexts
What they're looking for
- CUDA programming and optimization
- GPU microarchitecture knowledge (Ampere, Blackwell)
- C++ in large-scale codebases
- GPU performance profiling and debugging tools
- Linux development environments
- ML model optimization for inference
- Multiprocess systems debugging
- Motion planning or robotics domain experience
Benefits
- Health insurance
- Paid time off (vacation, sick leave, bereavement)
- Amazon Restricted Stock Units (RSUs)
- Zoox Stock Appreciation Rights
- Long-term and short-term disability insurance
- Life and long-term care insurance
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Zoox
Zoox develops autonomous vehicle technology and robotaxi systems, supported by manufacturing operations and AI validation infrastructure. The company is hiring for part-time student roles in hardware-software integration, manufacturing software engineering, AI testing and evaluation, and QA automation, as well as experienced engineers for simulation and AI performance assessment.
- Website
- zoox.com
Likely interview questions
- Describe your experience optimizing CUDA kernels for specific GPU architectures—what tools and techniques have you found most effective?
- Walk us through a time you identified and resolved a GPU performance bottleneck in a complex system. What was your approach?