xAI
Software Engineer - Platform Infrastructure (Rust, C++)
Palo Alto, CA$180k–$440kmidAdded today
About this role
Build and optimize large-scale distributed systems powering a massive supercomputing cluster for AI training. You'll work across the full stack—from GPU and Linux kernel optimization to container orchestration—in a small, high-performing team focused on systems excellence.
What you'll do
- Design and implement distributed systems for world-scale supercomputing clusters
- Profile, debug, and optimize performance across GPUs, Linux kernel, networking, and filesystems
- Collaborate on hardware-software-algorithm co-design to advance AI training capabilities
- Maintain and scale codebase for reliability and performance
- Develop internal tools to boost team productivity
What they're looking for
- C++, C, or Rust systems programming
- Computer systems fundamentals (transistors to applications)
- Kubernetes cluster architecture and operations
- Operating systems internals and Linux kernel concepts
- Performance profiling and low-level optimization
- Distributed systems debugging (perf, gdb, strace)
- Container technologies (Docker, containerd)
- Observability and monitoring (Prometheus, Grafana, OpenTelemetry)
Benefits
- Equity compensation
- Comprehensive medical, vision, and dental coverage
- 401(k) retirement plan
- Short and long-term disability insurance
- Life insurance
- Various discounts and perks
Opens the official application on the employer’s site. No login required.
xAI
xAI builds advanced AI infrastructure and systems, including the Grok model inference platform and Colossus GPU cluster. The company is hiring Mechanical, Electrical, and Facilities Engineers to design and maintain its data center operations, as well as Software Engineers to optimize high-performance inference systems and datacenter networking.
View all jobs at xAILikely interview questions
- Walk us through your experience optimizing performance in distributed systems—what specific bottlenecks did you identify and resolve?
- How have you debugged issues across multiple stack layers, and what tools did you use?