Skip to main content

Crusoe

Applied AI Inference Engineer

San Francisco, CA - USfulltimemidAdded 3 days ago

About this role

Crusoe seeks an Applied AI Inference Engineer to optimize large language models for production deployment, focusing on speed, cost, and reliability. You'll own the inference stack end-to-end—from profiling and kernel-level optimization to working directly with customers—turning performance gains into real-world impact across demanding ML workloads.

What you'll do

  • Profile and optimize LLM serving architectures, including prefill/decode disaggregation and request routing
  • Debug performance issues across the inference stack, from frameworks like vLLM and SGLang down to CUDA kernels
  • Tailor and scale deployments to meet customer-specific latency, throughput, and cost targets
  • Partner with customer engineering teams to move proofs of concept into monitored production services
  • Build software features and tooling around the inference stack using Python and general-purpose languages
  • Experiment rapidly, validate approaches, and ship well-tested optimizations with clear ownership end-to-end

What they're looking for

  • Python and/or C++ production experience
  • LLM inference optimization techniques
  • vLLM or SGLang framework proficiency
  • CUDA kernel profiling and analysis
  • Performance tuning and benchmarking
  • Distributed systems and serving architecture design
  • Customer collaboration and requirements gathering
  • Systems-level debugging and troubleshooting
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Crusoe

Crusoe builds AI infrastructure and data center systems, including cloud platforms, modular facilities, and manufacturing operations. The company is hiring engineers across software, mechanical engineering, facilities management, instrumentation, and CNC programming to design, optimize, and operate its mission-critical infrastructure.

View all jobs at Crusoe

Likely interview questions

  • Walk us through a time you optimized a production system for latency or throughput—what tools did you use to identify bottlenecks?
  • Describe your experience with LLM serving frameworks like vLLM or SGLang. How have you profiled or modified them?