Skip to main content

Adaption Labs

Inference Performance Engineer

San Francisco (Remote)fulltimemidAdded today

About this role

Own the performance and cost efficiency of an inference serving stack by optimizing caching, batching, quantization, and kernel-level operations. You'll work directly with the fleet operations team to improve throughput and latency across production workloads while maintaining reliability and model quality.

What you'll do

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization
  • Optimize long-context prefill and decode workloads based on production traffic patterns
  • Tune routing decisions between internal infrastructure and external providers for cost and performance trade-offs
  • Develop and maintain profiling systems to measure where time, memory, and compute are consumed
  • Work within and optimize serving engines like vLLM, SGLang, and TensorRT-LLM, diving into framework internals when needed
  • Collaborate with fleet operations engineers to validate and deploy performance improvements

What they're looking for

  • Model serving and inference infrastructure
  • Prefill and decode optimization, batching, and concurrency patterns
  • Python and at least one systems language (C++, Rust, or equivalent)
  • GPU performance profiling and optimization (CUDA, NCCL, memory layout)
  • Quantization and mixed-precision techniques
  • vLLM, SGLang, or TensorRT-LLM
  • Production performance engineering and cost optimization
  • Profiling and measurement tool development

Benefits

  • Flexible work arrangement with in-person collaboration in Bay Area and global team offsites
  • Annual travel stipend (Adaption Passport) to visit new countries
  • Weekly meal allowance for food delivery or groceries
  • Comprehensive medical benefits
  • Generous paid time off
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Adaption Labs

Adaption Labs builds machine learning systems and infrastructure that solve real-world customer problems at scale, from efficient AI inference to production ML deployments. The company is hiring Applied ML Engineers, distributed systems engineers, and research-focused technologists to develop adaptive AI solutions and bridge the gap between experimental research and reliable, deployed systems.

View all jobs at Adaption Labs

Likely interview questions

  • Walk us through a specific inference optimization project where you measurably reduced latency or cost—what was the bottleneck and how did you identify it?
  • How would you approach optimizing a long-context prefill workload that's currently memory-bound?