Sciforium
Distributed Training and Inference Engineer
- Confirmed live in the last 24 hours
- $190k–$250k
- Mid level
- Full-time
- On-site · San Francisco, CA
- Added 1 month ago
About this role
Sciforium, an AI infrastructure company backed by AMD, is seeking a Distributed Training and Inference Engineer to optimize and maintain their ML software stack. You'll be instrumental in building and improving the entire system, from low-level runtimes to high-level frameworks, ensuring fast, scalable, and efficient training and serving of large-scale AI models.
What you'll do
- Maintaining and optimizing ML libraries and frameworks (JAX, PyTorch, CUDA, ROCm).
- Building and improving the entire ML software stack.
- Ensuring efficient sharding and partitioning for distributed training and serving.
- Integrating and validating modules for runtime correctness and scalability.
- Conducting performance profiling and identifying bottlenecks.
- Troubleshooting hardware-software interaction issues.
What they're looking for
- Python
- JAX
- PyTorch
- CUDA
- ROCm
- NCCL
- XLA
Benefits
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary
- Equity
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Sciforium
Likely interview questions
- Describe your experience with distributed training frameworks like DTensor or GSPMD.
- How have you used profiling tools (Nsight, ROCm Profiler, etc.) to identify and resolve performance bottlenecks in ML workloads?