Vastai
AI, HPC & GPU Infrastructure Support Engineer
- Confirmed live in the last 24 hours
- $90k–$150k
- Mid level
- Full-time
- On-site · Los Angeles
- Added 2 months ago
About this role
Vast.ai is seeking an AI, HPC & GPU Infrastructure Support Engineer to troubleshoot complex infrastructure issues and provide expert-level support. This full-time role based in Los Angeles will involve resolving escalated tickets, improving diagnostic tooling, and collaborating with engineering and support teams to prevent future incidents. Ideal candidates possess strong Linux, GPU, and networking skills, and a passion for debugging and root cause analysis.
What you'll do
- Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments.
- Investigate GPU utilization, container resource constraints, and network-layer issues.
- Handle escalated support tickets involving GPU workload failures and container issues.
- Write and maintain internal runbooks and knowledge base articles.
- Build diagnostic and automation tooling in Python and Bash.
What they're looking for
- Linux (Ubuntu, RHEL/CentOS, Debian)
- Docker
- KVM/Virtualization
- NVIDIA GPU drivers and CUDA
- Python scripting
- Bash scripting
- Networking (VLAN, DNS, DHCP, VPN)
- Troubleshooting
Benefits
- Equity
- On-site location
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Vastai
Likely interview questions
- Describe a time you had to troubleshoot a complex GPU workload failure. What was your approach?
- How would you diagnose a high CPU utilization issue within a Docker container?