Elorian
Inference Infrastructure Engineer, Serving
- Confirmed live in the last 24 hours
- $200k–$400k
- Mid level
- Full-time
- Remote · Palo Alto
- 3+ yrs exp
- Added 2 months ago
About this role
Elorian AI is seeking an Inference Infrastructure Engineer to build and optimize the systems that serve their large multimodal AI models. This role focuses on achieving high-performance inference, ensuring smooth deployments, and supporting both research and real-world applications. The ideal candidate will be passionate about scaling ML services and reducing GPU costs.
What you'll do
- Build low-latency, high-throughput inference serving systems
- Implement optimization techniques like quantization and batching
- Optimize code and GPU utilization
- Implement multi-GPU/multi-node model parallelism
- Build autoscaling and load balancing
- Establish reliability and observability standards
What they're looking for
- Inference Serving Systems
- Quantization
- Batching
- Speculative Decoding
- KV Cache Management
- CUDA
- Python
- vLLM
- TensorRT-LLM
- Triton
- SGLang
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Equity
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Elorian
Likely interview questions
- Describe your experience building low-latency inference systems. What were the biggest challenges and how did you overcome them?
- Explain your understanding of quantization and its impact on model performance.