Build AI
ML Engineer, Inference Optimization
- Confirmed live in the last 24 hours
- $200k–$320k
- Mid level
- Full-time
- On-site · San Francisco
- Added 1 month ago
About this role
Build AI is seeking an ML Engineer to optimize inference performance and dramatically reduce compute costs. This role focuses on improving latency, throughput, and cost-effectiveness of models, working closely with research and product teams. The ideal candidate will be passionate about cost-driven optimization and scaling AI models efficiently.
What you'll do
- Optimize inference performance (latency, throughput, cost)
- Implement techniques like kernel optimization, batching, and quantization
- Profile and identify bottlenecks in inference pipelines
- Collaborate with research and product on cost-effective models
- Build and maintain serving and evaluation paths
- Establish cost as a primary metric
What they're looking for
- ML/Systems Engineering
- Inference Optimization
- Python
- Rust
- PyTorch
- Profiling Tools (e.g., Nsight, PyTorch Profiler)
- CUDA
Benefits
- Competitive pay
- Medical, dental, and vision
- Housing subsidy ($2k/month)
- Relocation support
- Wellness benefits
- Daily lunch and dinner
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Build AI
Likely interview questions
- Describe a time you significantly improved inference performance. What techniques did you use and what were the results?
- How do you approach profiling a complex ML pipeline to identify bottlenecks?