Skip to main content

Lm Studio

Software Engineer, Inference Runtime

  • Confirmed live in the last 24 hours
  • $150k–$350k
  • Mid level
  • Full-time
  • Remote · New York City
  • Added 2 months ago

About this role

LM Studio is seeking a Software Engineer to advance their inference runtime, both on-device and in the cloud. This role focuses on integrating new inference engines, optimizing model execution across various hardware, and contributing to open-source projects. The ideal candidate is passionate about human-AI interactions and thrives in a technically intense, collaborative environment.

What you'll do

  • Maintain and improve the inference stack
  • Integrate new model architectures
  • Optimize performance across various hardware
  • Build runtime capabilities (loading, batching, caching)
  • Benchmark and diagnose performance issues
  • Contribute to open-source projects

What they're looking for

  • Python
  • Transformer architectures
  • PyTorch
  • llama.cpp
  • MLX
  • ExecuTorch
  • vLLM

Benefits

  • Competitive salary and equity
  • Medical, vision, and dental benefits
  • Flexible PTO
  • Flexible WFH
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Lm Studio

Industry
Technology & Software
View all jobs at Lm Studio

Likely interview questions

  • Describe your experience building and optimizing production ML systems or inference runtimes.
  • Explain your understanding of transformer architectures and the inference process.