Lm Studio
Software Engineer, Inference Runtime
- Confirmed live in the last 24 hours
- $150k–$350k
- Mid level
- Full-time
- Remote · New York City
- Added 2 months ago
About this role
LM Studio is seeking a Software Engineer to advance their inference runtime, both on-device and in the cloud. This role focuses on integrating new inference engines, optimizing model execution across various hardware, and contributing to open-source projects. The ideal candidate is passionate about human-AI interactions and thrives in a technically intense, collaborative environment.
What you'll do
- Maintain and improve the inference stack
- Integrate new model architectures
- Optimize performance across various hardware
- Build runtime capabilities (loading, batching, caching)
- Benchmark and diagnose performance issues
- Contribute to open-source projects
What they're looking for
- Python
- Transformer architectures
- PyTorch
- llama.cpp
- MLX
- ExecuTorch
- vLLM
Benefits
- Competitive salary and equity
- Medical, vision, and dental benefits
- Flexible PTO
- Flexible WFH
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lm Studio
- Industry
- Technology & Software
Likely interview questions
- Describe your experience building and optimizing production ML systems or inference runtimes.
- Explain your understanding of transformer architectures and the inference process.