Mistral AI
Research Engineer, ML Platform
About this role
Mistral AI seeks a Research Platform Engineer to design and operate ML infrastructure for large-scale distributed training, evaluation, and inference. You'll build orchestration systems, manage GPU workloads across clusters, and create self-service tools that enable researchers to run complex distributed workloads reliably.
What you'll do
- Design and develop ML platform services, APIs, and controllers for training, fine-tuning, evaluation, and batch inference
- Build workload orchestration systems including queueing, scheduling, admission control, quotas, and topology-aware placement
- Manage provisioning and allocation of heterogeneous GPU resources across multi-region clusters
- Create self-service workflows and debugging tools to improve researcher experience with distributed workloads
- Optimize GPU utilization, scheduling latency, startup time, and overall infrastructure efficiency
- Participate in on-call rotations and troubleshoot production issues across orchestration, networking, storage, and GPU infrastructure
What they're looking for
- Kubernetes platform engineering (controllers, operators, CRDs, scheduling, networking, storage)
- Python or Go programming in production systems
- GPU infrastructure and technologies (CUDA, NCCL, PyTorch, high-performance networking)
- ML workload orchestration tools (Kueue, Karpenter, Volcano, Kyverno)
- Distributed systems design and troubleshooting
- Workload scheduling concepts (quotas, priorities, preemption, gang scheduling)
- Performance and reliability diagnostics across software, hardware, and infrastructure layers
- Developer experience and API design
Benefits
- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal and transportation allowances
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mistral AI
Mistral AI builds large-scale machine learning systems and open-weight AI models, supported by infrastructure powering petabyte-scale HPC clusters and enterprise AI platforms. The company is hiring Systems Engineers, Research Engineers, Site Reliability Engineers, and Applied AI Engineers to scale training infrastructure, optimize data systems, ensure platform reliability, and drive customer adoption across industries.
- Website
- mistral.ai
Likely interview questions
- Describe your experience building or operating Kubernetes-based ML platforms. What were the key scaling challenges and how did you address them?
- How would you approach scheduling and prioritizing GPU workloads across heterogeneous hardware in a multi-cluster environment?