Skip to main content

Andromeda

Forward Deployed Engineer - SRE

North America Remote / San Francisco, CA (Remote)fulltimemidAdded today

About this role

Andromeda is hiring a Forward Deployed Engineer to embed with customers running large-scale AI training and inference workloads, diagnosing performance issues across the full stack—from GPU fabric to frameworks—while building automation and reliability into the platform.

What you'll do

  • Serve as primary technical contact for customers running distributed training/inference, owning full onboarding and ongoing engagement
  • Diagnose and resolve multi-layer failures in GPU clusters including NCCL timeouts, stragglers, I/O stalls, and driver mismatches
  • Profile and optimize distributed training performance to improve MFU, reduce idle GPU time, and accelerate time-to-first-run
  • Monitor and maintain health of high-speed interconnects (InfiniBand, RoCE, NVLink) and diagnose fabric-level issues
  • Automate recurring deployment problems: cluster provisioning, GPU burn-in, preflight checks, firmware lifecycle, and reference configurations
  • Lead incident response for complex failures spanning hardware, networking, orchestration, and ML frameworks

What they're looking for

  • GPU cluster operations (NVIDIA A100/H100/H200)
  • High-speed fabric networking (InfiniBand, RoCE, NVLink)
  • Linux kernel tuning and driver management
  • Distributed training systems and performance profiling
  • Container runtimes and orchestration (Kubernetes, Slurm)
  • CUDA toolkit and GPU memory management
  • Incident response and root cause analysis
  • Infrastructure automation and monitoring
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Andromeda

Andromeda is an AI infrastructure startup that builds hardware and software platforms for large-scale AI training and inference, operating a compute marketplace that routes jobs across multiple cloud providers. The company is hiring Solutions Engineers, Site Reliability Engineers, and Infrastructure Platform Engineers to deliver customer solutions, manage Kubernetes clusters, and design core orchestration systems.

View all jobs at Andromeda

Likely interview questions

  • Walk us through a time you debugged a production GPU cluster issue—what was failing and how did you isolate the root cause?
  • Describe your experience with InfiniBand or RoCE fabrics in distributed training. How have you diagnosed a degraded link or congestion?