ElevenLabs
HPC Infrastructure Engineer - GPU Clusters
About this role
ElevenLabs seeks an HPC Infrastructure Engineer to design, deploy, and maintain GPU clusters supporting AI model training and inference at scale. You'll own the full lifecycle of high-performance computing infrastructure for a rapidly growing AI audio company backed by top-tier investors.
What you'll do
- Design and provision GPU cluster architectures for training and inference workloads
- Deploy, configure, and optimize CUDA, container orchestration, and distributed computing frameworks
- Monitor cluster performance, troubleshoot hardware/software issues, and implement auto-scaling solutions
- Collaborate with ML researchers to optimize resource allocation and job scheduling
- Implement infrastructure-as-code practices and maintain cluster documentation
- Manage cost optimization and capacity planning across multi-GPU environments
What they're looking for
- GPU cluster management (NVIDIA CUDA, cuDNN, cluster schedulers)
- Container orchestration (Kubernetes, Docker)
- Linux system administration and networking
- Infrastructure-as-code (Terraform, Ansible)
- Distributed systems and HPC concepts
- Python or Go for automation scripting
- Cloud platforms (AWS, GCP, or similar)
- Monitoring and observability tools (Prometheus, Grafana)
Benefits
- Annual professional development stipend
- Annual social travel stipend for team meetups
- Monthly co-working allowance for remote work
- Company offsites in international locations
- Opportunity to work on generational AI technology
- Global team with flexible location policy
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ElevenLabs
ElevenLabs builds an AI voice platform that enterprises and developers integrate into their applications and workflows. The company is hiring for compliance, developer advocacy, customer-facing engineering, and enterprise sales engineering roles to support platform adoption across regulated industries and strategic customer deployments.
- Website
- elevenlabs.io
Likely interview questions
- Describe your experience designing and scaling GPU clusters—what challenges did you encounter and how did you resolve them?
- How do you approach performance optimization for distributed training workloads across multiple GPUs?