Skip to main content

Pluralis Research

Research Engineer - Pre-training

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • Remote · San Francisco
  • Added 1 month ago

About this role

Pluralis Research is seeking a Research Engineer to build a distributed training system capable of scaling Protocol Learning to frontier model sizes. This role focuses on implementing and optimizing model-parallel training across geographically dispersed, heterogeneous hardware over ordinary internet connections, addressing challenges like fault tolerance and communication efficiency. You'll be instrumental in shaping the future of decentralized AI training.

What you'll do

  • Implement and optimize model-parallel training for large models.
  • Reduce communication overhead while maintaining model convergence.
  • Ensure runs survive node churn through robust checkpointing and state synchronization.
  • Build monitoring tools to track throughput, bottlenecks, and model quality.

What they're looking for

  • Distributed training
  • PyTorch
  • FSDP
  • DeepSpeed
  • Megatron
  • Python
  • Concurrency
  • Profiling

Benefits

  • Equity-Heavy Package
  • Remote-First Culture
  • Visa Sponsorship (Australia or US)
  • Flexible work environment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Pluralis Research

View all jobs at Pluralis Research

Likely interview questions

  • Describe your experience with distributed training frameworks like FSDP, DeepSpeed, or Megatron.
  • How have you approached performance optimization in distributed systems, specifically addressing communication overhead?