Skip to main content

Elorian

Inference Infrastructure Engineer, Serving

  • Confirmed live in the last 24 hours
  • $200k–$400k
  • Mid level
  • Full-time
  • Remote · Palo Alto
  • 3+ yrs exp
  • Added 2 months ago

About this role

Elorian AI is seeking an Inference Infrastructure Engineer to build and optimize the systems that serve their large multimodal AI models. This role focuses on achieving high-performance inference, ensuring smooth deployments, and supporting both research and real-world applications. The ideal candidate will be passionate about scaling ML services and reducing GPU costs.

What you'll do

  • Build low-latency, high-throughput inference serving systems
  • Implement optimization techniques like quantization and batching
  • Optimize code and GPU utilization
  • Implement multi-GPU/multi-node model parallelism
  • Build autoscaling and load balancing
  • Establish reliability and observability standards

What they're looking for

  • Inference Serving Systems
  • Quantization
  • Batching
  • Speculative Decoding
  • KV Cache Management
  • CUDA
  • Python
  • vLLM
  • TensorRT-LLM
  • Triton
  • SGLang

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Equity
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Elorian

View all jobs at Elorian

Likely interview questions

  • Describe your experience building low-latency inference systems. What were the biggest challenges and how did you overcome them?
  • Explain your understanding of quantization and its impact on model performance.