Skip to main content

Figure

Helix AI Engineer, Training Performance

San Jose, CA$200k–$400kfull-timemidAdded today

About this role

Figure seeks an AI Training Performance Engineer to optimize distributed training for 100B+ parameter models across massive GPU clusters. You'll focus on kernel optimization, performance monitoring, hardware co-design, and evaluating emerging accelerators to maximize training efficiency at scale.

What you'll do

  • Optimize training performance for 100B+ parameter models across 100k+ GPUs
  • Write and optimize custom kernels in Triton/CUDA and contribute to kernel compilers
  • Build monitoring dashboards and tools for performance regression detection and root-cause analysis
  • Optimize data loading pipelines and implement checkpointing/fault tolerance strategies
  • Collaborate on accelerator selection, cluster topology, and hardware procurement decisions
  • Evaluate emerging accelerators (AMD, TPU, ASICs) and lead proof-of-concept benchmarks

What they're looking for

  • GPU architecture and performance profiling (Nsight, PyTorch Profiler)
  • Python and CUDA/C++ development
  • Collective communication protocols (NCCL, RDMA, InfiniBand)
  • Distributed training frameworks (FSDP, model/data parallelism strategies)
  • Performance debugging at scale (stragglers, OOMs, numerical divergence)
  • Hardware efficiency metrics (MFU/HFU) and optimization reasoning
  • Kernel optimization and compiler extensions (Triton, Gluon)
  • Multi-accelerator and heterogeneous system evaluation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Figure

Figure develops advanced humanoid robots powered by AI technology. The company is hiring engineers across mechanical design, firmware development, manufacturing, quality assurance, and security to build and refine its autonomous robotic systems.

View all jobs at Figure

Likely interview questions

  • Walk us through a large-scale performance optimization project you led—what was the bottleneck, how did you identify it, and what was the impact?
  • How would you approach optimizing a memory-bound operation on a GPU, and what profiling tools would you use?