Vastai
Systems Operations Support Engineer — Linux
- Confirmed live in the last 24 hours
- $90k–$150k
- Mid level
- Full-time
- On-site · Los Angeles
- Added 2 months ago
About this role
Vast.ai is seeking a Systems Operations Support Engineer to tackle complex infrastructure issues and contribute to their rapidly growing AI cloud platform. This full-time role, based in Los Angeles, involves deep troubleshooting across hardware, networking, and containerized environments, with a focus on GPU workloads. You'll be a key engineering resource for the L1 support team, proactively identifying patterns and building tools to prevent future incidents.
What you'll do
- Handle escalated support tickets related to GPU workloads, containers, and networking.
- Assist with supplier onboarding and machine management, providing technical guidance.
- Troubleshoot issues across Docker, NVIDIA CUDA/GPU drivers, and KVM virtualization.
- Develop and maintain internal runbooks and diagnostic tooling using Python and Bash.
- Collaborate with engineering and support teams to address systemic platform issues.
What they're looking for
- Linux (Ubuntu, RHEL/CentOS, Debian)
- Docker
- NVIDIA CUDA/GPU
- KVM Virtualization
- Networking (VLAN, DNS, DHCP)
- Python
- Bash
- Troubleshooting
Benefits
- Equity
- Globally distributed team
- Opportunity to work with AI systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Vastai
Likely interview questions
- Describe a time you diagnosed and resolved a complex infrastructure issue. What was your approach?
- How would you approach troubleshooting a performance bottleneck involving GPU utilization?