Skip to main content

Sciforium

Distributed Training and Inference Engineer

  • Confirmed live in the last 24 hours
  • $190k–$250k
  • Mid level
  • Full-time
  • On-site · San Francisco, CA
  • Added 1 month ago

About this role

Sciforium, an AI infrastructure company backed by AMD, is seeking a Distributed Training and Inference Engineer to optimize and maintain their ML software stack. You'll be instrumental in building and improving the entire system, from low-level runtimes to high-level frameworks, ensuring fast, scalable, and efficient training and serving of large-scale AI models.

What you'll do

  • Maintaining and optimizing ML libraries and frameworks (JAX, PyTorch, CUDA, ROCm).
  • Building and improving the entire ML software stack.
  • Ensuring efficient sharding and partitioning for distributed training and serving.
  • Integrating and validating modules for runtime correctness and scalability.
  • Conducting performance profiling and identifying bottlenecks.
  • Troubleshooting hardware-software interaction issues.

What they're looking for

  • Python
  • JAX
  • PyTorch
  • CUDA
  • ROCm
  • NCCL
  • XLA

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary
  • Equity
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Sciforium

View all jobs at Sciforium

Likely interview questions

  • Describe your experience with distributed training frameworks like DTensor or GSPMD.
  • How have you used profiling tools (Nsight, ROCm Profiler, etc.) to identify and resolve performance bottlenecks in ML workloads?