Sciforium
GPU Cluster Engineer, Systems & Platform
- Confirmed live in the last 24 hours
- $150k–$220k
- Mid level
- Full-time
- On-site · San Francisco, CA
- 5+ yrs exp
- Added 2 months ago
About this role
Sciforium is seeking a skilled GPU Cluster Engineer to build and maintain the software foundation for their rapidly scaling AI infrastructure. This role involves owning the entire GPU cluster software stack, from kernel tuning to ML frameworks, ensuring optimal performance for both model training and serving teams. The ideal candidate will automate node provisioning, manage fleet consistency, and troubleshoot complex performance issues.
What you'll do
- Automate node software definition, including OS images, kernel tuning, and driver stacks.
- Build automated acceptance suites to validate node performance and stability.
- Manage and maintain GPU driver and CUDA/ROCm compatibility across the fleet.
- Automate infrastructure configuration using Ansible or SaltStack.
- Deploy and operate GPU-enabled Kubernetes for inference workloads and Slurm for multi-node training.
- Develop Python/Bash tooling for cluster operations, health reporting, and workflow automation.
What they're looking for
- Linux Internals
- NVIDIA (CUDA)
- AMD (ROCm)
- Kubernetes
- Slurm/Run:AI
- Ansible/SaltStack
- Python
- Bash
- NCCL/RCCL
- RDMA
Benefits
- Backed by multi-million-dollar funding
- Direct sponsorship from AMD
- Rapidly scaling team and environment
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Sciforium
Likely interview questions
- Describe your experience with kernel tuning for GPU clusters, specifically mentioning NUMA or cgroups.
- How would you approach debugging an NCCL hang or timeout in a multi-node training environment?