Skip to main content

Lambda

HPC Support Engineer

Remote, USA (Remote)$122k–$162kfulltimemidAdded today

About this role

Lambda seeks a senior HPC Support Engineer to troubleshoot complex infrastructure and platform issues in their AI cloud, working down to hardware and kernel levels. You'll lead technical escalations, mentor peers, participate in on-call rotations, and collaborate with engineering teams to turn customer pain points into permanent fixes.

What you'll do

  • Serve as senior technical escalation point for infrastructure and platform issues, troubleshooting hardware, drivers, and kernel-level problems
  • Distinguish between hardware failures, driver issues, kernel problems, and customer misconfiguration to resolve issues correctly
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
  • Proactively identify process and tooling gaps, then fix them using AI-assisted scripting and automation
  • Lead on-call rotation, owning major incidents and critical customer issues
  • Mentor junior support engineers and collaborate with engineering teams on permanent solutions

What they're looking for

  • HPC administration and support (3+ years hands-on)
  • Linux system administration and cluster management
  • Kubernetes and/or Slurm for cluster orchestration
  • CUDA, NCCL, NVLink, and GPUDirect RDMA
  • Monitoring and logging tools (Prometheus, Grafana, Datadog)
  • Kernel debugging, log analysis, and performance profiling
  • High throughput networking (InfiniBand/RoCE)
  • CI/CD and coding with AI-assisted development
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Lambda

Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.

View all jobs at Lambda

Likely interview questions

  • Walk us through a time you debugged a kernel-level issue in an HPC cluster—what was your approach and how did you isolate the root cause?
  • Describe your experience with CUDA and GPU communication technologies like NCCL or GPUDirect RDMA. Have you troubleshot performance issues with these?