Skip to main content

Pluralis Research

Research Engineer - Post-Training

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • Remote · San Francisco
  • Added 1 month ago

About this role

Pluralis Research is seeking a Research Engineer to tackle the unique challenges of post-training large language models in a decentralized environment. This role involves building an end-to-end RL training stack for consumer-grade devices over the internet, adapting existing algorithms for asynchronous, high-latency scenarios, and ultimately deploying the first post-trained models. The ideal candidate will have hands-on experience with RL post-training, strong engineering skills, and a belief in the potential of Protocol Learning.

What you'll do

  • Build the RL training loop, including rollout ingestion, reward computation, and policy updates.
  • Adapt existing RL algorithms to handle asynchronous, high-latency, and partially trusted environments.
  • Develop evaluation metrics to track model improvement and deploy post-trained models.
  • Set the technical direction and drive projects to completion.
  • Ship the first decentralized post-trained release.

What they're looking for

  • RL Post-training (RLHF, RLVR)
  • Python
  • PyTorch
  • Concurrency
  • Failure Handling
  • Profiling
  • Distributed RL

Benefits

  • Equity-Heavy Package
  • Remote-First Culture
  • Visa Sponsorship (Australia or US)
  • Flexible Work Environment
  • Open Problems
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Pluralis Research

View all jobs at Pluralis Research

Likely interview questions

  • Describe your experience with RL post-training on large language models, including the systems you've worked with.
  • How would you approach adapting standard RL algorithms to handle high-latency, asynchronous rollouts?