Chai Discovery
Research Engineer - ML Infrastructure
About this role
Chai Discovery seeks a Research Engineer to optimize and scale the distributed ML training infrastructure that powers their molecular design platform. You'll architect high-performance training systems, debug bottlenecks across GPU clusters, and collaborate with research scientists to ensure frontier-scale model training runs reliably and efficiently.
What you'll do
- Architect, debug, and optimize distributed ML training stacks across model, layer, and kernel levels
- Profile end-to-end training runs to identify bottlenecks in compute, communication, and storage
- Build monitoring tooling to track throughput, utilization, and uptime across GPU clusters
- Optimize ML workloads using parallelism strategies, quantization, and custom CUDA/Triton kernels
- Partner with Research Scientists to ensure new model architectures scale efficiently from experiments to production
- Own reliability and fault tolerance of the training stack, including checkpointing and deterministic orchestration
What they're looking for
- Python
- PyTorch or JAX
- Distributed systems design
- GPU cluster orchestration
- CUDA/Triton kernel optimization
- ML workload optimization
- Systems-level debugging and profiling
- Fault tolerance and checkpointing systems
Benefits
- Work on frontier AI research in drug discovery and molecular design
- High-velocity, ownership-driven culture
- Competitive compensation
- Opportunity to work with world-class research teams
- San Francisco office location
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Chai Discovery
Chai Discovery builds AI-powered tools for molecular biology and drug discovery, enabling scientists to design new therapeutic molecules through advanced AI systems. The company is hiring Security Engineers, Design Engineers, AI Research Engineers, Infrastructure Engineers, and Product Software Engineers to develop and scale its platform.
View all jobs at Chai DiscoveryLikely interview questions
- Walk us through a time you identified and resolved a critical bottleneck in a distributed training system—what was your debugging approach?
- How do you decide between different parallelism strategies (data, tensor, pipeline) for a given model and cluster setup?