Mistral AI
Applied Scientist/Research Engineer - Palo Alto/NYC
About this role
Mistral AI seeks Applied Scientists and Research Engineers to develop state-of-the-art AI models across multiple modalities and deploy them at scale on GPU clusters. You'll drive innovative research, collaborate with enterprise clients on complex use cases, and build the tools and frameworks that power next-generation AI systems.
What you'll do
- Train and deploy large-scale models using distributed GPU clusters, handling infrastructure challenges like memory optimization and network communication
- Generate, curate, and evaluate datasets for pre-training and post-training pipelines to meet performance targets
- Build tools, frameworks, and pipelines for data processing, model training, evaluation, and deployment
- Design and implement agent-based and RAG solutions for diverse enterprise use cases
- Manage research projects and maintain technical communication with external client research teams
- Contribute to large codebases with clean, high-performance, fault-tolerant Python implementations
What they're looking for
- PyTorch or JAX
- Python (advanced)
- Distributed training and GPU optimization
- Machine learning model development
- Data curation and evaluation
- Agents and RAG systems
- Large codebase navigation
- Cross-functional collaboration and communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Mistral AI
Mistral AI builds large-scale machine learning systems and open-weight AI models, supported by infrastructure powering petabyte-scale HPC clusters and enterprise AI platforms. The company is hiring Systems Engineers, Research Engineers, Site Reliability Engineers, and Applied AI Engineers to scale training infrastructure, optimize data systems, ensure platform reliability, and drive customer adoption across industries.
- Website
- mistral.ai
Likely interview questions
- Walk us through your experience training models on distributed GPU clusters—what scaling challenges have you faced and how did you solve them?
- Describe a research project where you developed a novel method and then applied it to a real-world use case. What was the impact?