Mirage
Research Engineer, Generative Video
About this role
Mirage is hiring a Research Engineer to build and optimize large-scale video generation models, focusing on training infrastructure, inference efficiency, and productionizing cutting-edge generative video systems. You'll work on novel modeling approaches, scaling strategies, and optimization techniques to enable real-time, low-latency video generation at production scale.
What you'll do
- Train and optimize large-scale video and multimodal models
- Improve efficiency across training and inference using distillation, quantization, and pruning
- Build and maintain distributed training systems with optimized GPU utilization and parallelism
- Develop tooling for experimentation, evaluation, and debugging of generative models
- Translate research models into robust, production-ready systems
- Monitor and improve model performance in real-world usage
What they're looking for
- Deep learning systems and infrastructure
- PyTorch and CUDA programming
- Triton and distributed training (FSDP)
- Model optimization (distillation, quantization, pruning)
- Low-latency inference optimization
- Performance profiling and debugging
- Diffusion and autoregressive model architectures
- GPU utilization and throughput optimization
Benefits
- Medical, dental, and vision insurance
- 401K with employer match
- Commuter benefits
- Catered lunches multiple days per week
- Generous PTO policy
- Team offsites and monthly events
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mirage
Mirage builds an AI-native video platform that leverages generative media and large language models to enable sophisticated video production, editing, and creative workflows. The company is hiring backend engineers, full-stack software engineers, ML engineers, and iOS developers to advance their AI-driven platform and enhance user experiences in web-based and mobile media creation.
View all jobs at MirageLikely interview questions
- Walk us through your experience optimizing large models for low-latency inference—what techniques did you use and what was the impact?
- Describe a time you scaled a distributed training system; what bottlenecks did you encounter and how did you resolve them?