Skip to main content

Sciforium

GPU Cluster Engineer, Systems & Platform

  • Confirmed live in the last 24 hours
  • $150k–$220k
  • Mid level
  • Full-time
  • On-site · San Francisco, CA
  • 5+ yrs exp
  • Added 2 months ago

About this role

Sciforium is seeking a skilled GPU Cluster Engineer to build and maintain the software foundation for their rapidly scaling AI infrastructure. This role involves owning the entire GPU cluster software stack, from kernel tuning to ML frameworks, ensuring optimal performance for both model training and serving teams. The ideal candidate will automate node provisioning, manage fleet consistency, and troubleshoot complex performance issues.

What you'll do

  • Automate node software definition, including OS images, kernel tuning, and driver stacks.
  • Build automated acceptance suites to validate node performance and stability.
  • Manage and maintain GPU driver and CUDA/ROCm compatibility across the fleet.
  • Automate infrastructure configuration using Ansible or SaltStack.
  • Deploy and operate GPU-enabled Kubernetes for inference workloads and Slurm for multi-node training.
  • Develop Python/Bash tooling for cluster operations, health reporting, and workflow automation.

What they're looking for

  • Linux Internals
  • NVIDIA (CUDA)
  • AMD (ROCm)
  • Kubernetes
  • Slurm/Run:AI
  • Ansible/SaltStack
  • Python
  • Bash
  • NCCL/RCCL
  • RDMA

Benefits

  • Backed by multi-million-dollar funding
  • Direct sponsorship from AMD
  • Rapidly scaling team and environment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Sciforium

View all jobs at Sciforium

Likely interview questions

  • Describe your experience with kernel tuning for GPU clusters, specifically mentioning NUMA or cgroups.
  • How would you approach debugging an NCCL hang or timeout in a multi-node training environment?