Skip to main content

Applied Intuition

Machine Learning Performance Engineer - Offboard Training & Inference

Sunnyvale$215k–$285kfulltimemidAdded yesterday

About this role

Applied Intuition seeks a Machine Learning Performance Engineer to optimize distributed training and large-scale batch inference workloads in the datacenter. The role focuses on profiling and improving throughput, cluster efficiency, and cost-per-unit-data across accelerator infrastructure, data pipelines, and ML frameworks.

What you'll do

  • Profile and optimize end-to-end distributed training including data loading, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize petabyte-scale offline batch inference through batching strategies, quantization, graph optimization, and accelerator saturation
  • Establish performance models and roofline analysis to quantify gaps between theoretical and achieved performance
  • Improve multi-node scaling efficiency by analyzing sharding, parallelism, collective communication, and memory bandwidth bottlenecks
  • Reduce GPU idle time and improve cluster goodput by addressing I/O stalls, scheduling gaps, and failure recovery on long-running jobs
  • Build benchmarking, observability, and regression-detection tooling to prevent silent performance degradation

What they're looking for

  • ML performance engineering and profiling
  • Distributed multi-node training frameworks (FSDP, DeepSpeed, Megatron, NCCL)
  • GPU and accelerator performance optimization
  • High-throughput batch inference systems (Triton, TensorRT, ONNX Runtime, Ray)
  • Python and C++ or systems language
  • Roofline analysis and performance modeling
  • Debugging and root-cause investigation
  • Machine learning fundamentals
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Applied Intuition

Applied Intuition builds autonomous vehicle and defense systems software, including motion planning algorithms, simulation infrastructure, and autonomy integration platforms for aerial and ground platforms. The company is hiring for security engineers, robotics/autonomy software engineers, hardware-in-the-loop specialists, and IT operations professionals to support its growing physical AI operations.

View all jobs at Applied Intuition

Likely interview questions

  • Walk us through a time you identified and fixed a performance bottleneck in a distributed training system—what tools did you use and what was the impact?
  • How would you approach profiling and optimizing a multi-node training job that shows poor scaling efficiency as you increase node count?