xAI
ML Infrastructure Engineer
About this role
SpaceXAI seeks an ML Infrastructure Engineer to design and optimize the GPU compute infrastructure, training frameworks, and data pipelines supporting demanding machine learning workloads. You'll work with ML teams to ensure scalability and reliability across the full stack in a small, high-performing organization.
What you'll do
- Design, build, and scale GPU compute infrastructure and training frameworks for rapid ML experimentation
- Develop and integrate large-scale data pipelines, training systems, and inference infrastructure
- Collaborate with ML teams to productionize models and ensure seamless stack integration
- Optimize scalability, reliability, and efficiency of production machine learning systems
- Solve complex technical problems independently across the full stack
- Mentor junior engineers and contribute to team growth
What they're looking for
- Python programming
- C++ or Rust
- JAX or PyTorch frameworks
- Distributed systems and GPU infrastructure
- CUDA, NVIDIA drivers, and networking fundamentals
- Linux systems and container orchestration
- Job schedulers like Slurm
- Infrastructure tooling and configuration management
Benefits
- Equity compensation
- Comprehensive medical, vision, and dental coverage
- 401(k) retirement plan
- Short and long-term disability insurance
- Life insurance
- Various discounts and perks
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
xAI
xAI builds advanced AI infrastructure and systems, including the Grok model inference platform and Colossus GPU cluster. The company is hiring Mechanical, Electrical, and Facilities Engineers to design and maintain its data center operations, as well as Software Engineers to optimize high-performance inference systems and datacenter networking.
View all jobs at xAILikely interview questions
- Describe your experience designing and scaling GPU infrastructure for production ML workloads—what challenges did you face?
- Walk us through how you've built or optimized data pipelines for large-scale training systems.