Skip to main content

Build AI

ML Engineer, Inference Optimization

  • Confirmed live in the last 24 hours
  • $200k–$320k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 1 month ago

About this role

Build AI is seeking an ML Engineer to optimize inference performance and dramatically reduce compute costs. This role focuses on improving latency, throughput, and cost-effectiveness of models, working closely with research and product teams. The ideal candidate will be passionate about cost-driven optimization and scaling AI models efficiently.

What you'll do

  • Optimize inference performance (latency, throughput, cost)
  • Implement techniques like kernel optimization, batching, and quantization
  • Profile and identify bottlenecks in inference pipelines
  • Collaborate with research and product on cost-effective models
  • Build and maintain serving and evaluation paths
  • Establish cost as a primary metric

What they're looking for

  • ML/Systems Engineering
  • Inference Optimization
  • Python
  • Rust
  • PyTorch
  • Profiling Tools (e.g., Nsight, PyTorch Profiler)
  • CUDA

Benefits

  • Competitive pay
  • Medical, dental, and vision
  • Housing subsidy ($2k/month)
  • Relocation support
  • Wellness benefits
  • Daily lunch and dinner
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Build AI

View all jobs at Build AI

Likely interview questions

  • Describe a time you significantly improved inference performance. What techniques did you use and what were the results?
  • How do you approach profiling a complex ML pipeline to identify bottlenecks?