SpaceX
Software Engineer, Inference (AI Data Engineering)
About this role
SpaceX seeks a Software Engineer to design and optimize high-performance AI inference platforms that serve mission-critical models across the organization. You'll own distributed infrastructure, low-level GPU optimizations, and end-to-end serving systems supporting SpaceX's most ambitious engineering goals.
What you'll do
- Design and build highly reliable, high-throughput inference systems serving AI models across SpaceX
- Architect scalable distributed infrastructure including load balancing, auto-scaling, batch scheduling, and KV cache management
- Optimize inference latency and throughput via GPU kernels, quantization, speculative decoding, and other acceleration techniques
- Develop mission-critical serving systems with 100% uptime, low tail latency, and strong observability
- Benchmark and accelerate inference engines like SGLang, vLLM, and TensorRT-LLM for production workloads
- Own end-to-end components including request routing, SDK development, rate limiting, and CI/CD infrastructure
What they're looking for
- Distributed systems design and implementation
- Rust or C++ for systems programming
- GPU optimization and kernel programming
- LLM inference engines (vLLM, SGLang, TensorRT-LLM, Triton)
- High-concurrency production serving systems
- Backend development and full-stack engineering
- Service observability and reliability practices
- Quantization, batching, caching, and parallelism
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
SpaceX
SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.
- Website
- spacex.com
Likely interview questions
- Describe your experience designing or maintaining distributed systems at scale—what were the key reliability and performance challenges?
- Have you optimized inference latency or throughput in production? Walk us through your approach and specific techniques used.