Crusoe
Software Engineer II (DCIE)
About this role
Crusoe seeks a Software Engineer II to develop diagnostics, automation, and observability tooling for managing large-scale GPU server fleets and data center infrastructure. You'll own the design and deployment of solutions for hardware fault detection, remediation, and system validation across NVIDIA and AMD GPU platforms.
What you'll do
- Build diagnostic and troubleshooting tools for GPU racks and high-density compute systems
- Develop automation and AI agents for component-level hardware diagnosis and remediation
- Create post-repair validation and testing tooling using burn-in, PyTorch, and NVIDIA NCCL
- Collaborate with data center operations to build critical environment management tooling
- Deploy, monitor, and operationally support developed solutions to maximize GPU fleet availability
- Develop automation for facilities management including power and direct liquid cooling systems
What they're looking for
- Software engineering (2–3 years experience)
- Distributed systems and cloud platforms (Kubernetes, IaC, GCP)
- Programming in Go, Python, Java, or Rust
- Hardware diagnostics and troubleshooting
- GPU platform expertise (NVIDIA A100/H200/GB200/B200, AMD 350X/355X)
- Problem-solving and rapid prototyping
- Technical leadership and project execution
Benefits
- Competitive compensation
- Restricted Stock Units
- Health insurance (HDHP, PPO, vision, dental)
- HSA employer contributions and 401(k) 100% match up to 4%
- Paid parental leave, life insurance, and disability coverage
- Generous paid time off and cell phone reimbursement
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Crusoe
Crusoe builds AI infrastructure and data center systems, including cloud platforms, modular facilities, and manufacturing operations. The company is hiring engineers across software, mechanical engineering, facilities management, instrumentation, and CNC programming to design, optimize, and operate its mission-critical infrastructure.
View all jobs at CrusoeLikely interview questions
- Tell us about a time you identified a problem in a distributed system and rapidly shipped a scalable solution—what was your approach?
- Describe your experience with GPU platforms or hardware-level diagnostics. What challenges did you face?