xAI
Network Engineer
About this role
SpaceXAI seeks a hands-on Network Engineer to design, deploy, and operate production datacenter and high-performance computing networks at scale. You'll own routing and switching infrastructure, troubleshoot complex Layer 2/3 incidents, and drive network reliability as the company builds AI systems. This role involves travel to build sites, on-call participation, and direct collaboration with compute and infrastructure teams.
What you'll do
- Design, deploy, and operate production datacenter and core networks at scale with focus on availability and performance
- Own BGP and IGP routing configuration standards including change design, peer review, and safe execution
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
- Troubleshoot Layer 2/Layer 3 incidents end-to-end and drive root cause analysis and lasting fixes
- Build monitoring, alerting, and operational documentation to catch and resolve issues quickly
- Automate repetitive network tasks using Python, Ansible, or similar tooling to reduce operational burden
What they're looking for
- BGP and interior routing protocols (OSPF, IS-IS)
- TCP/IP, VLANs, EVPN/VXLAN, and high-speed Ethernet optics
- Production network troubleshooting and incident response
- Network automation (Python, Ansible, Terraform)
- Datacenter vendor platforms (Arista, Cisco, Juniper, Nvidia/Mellanox)
- High-performance/supercompute networking (RoCEv2, GPU cluster fabrics preferred)
- Leaf-spine architecture and large-scale Ethernet fabrics
- Clear technical communication and documentation
Benefits
- Competitive base salary with equity component
- Comprehensive medical, vision, and dental coverage
- 401(k) retirement plan access
- Short and long-term disability insurance
- Life insurance
- Various discounts and perks
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
xAI
xAI builds advanced AI infrastructure and systems, including the Grok model inference platform and Colossus GPU cluster. The company is hiring Mechanical, Electrical, and Facilities Engineers to design and maintain its data center operations, as well as Software Engineers to optimize high-performance inference systems and datacenter networking.
View all jobs at xAILikely interview questions
- Walk us through a complex Layer 2/3 network incident you troubleshot in production—what was your approach and how did you verify the fix?
- Describe your experience designing and operating BGP in a large-scale datacenter. How do you handle route propagation and failover scenarios?