Pluralis Research
Research Engineer - Post-Training
- Confirmed live in the last 24 hours
- No salary listed
- Mid level
- Full-time
- Remote · San Francisco
- Added 1 month ago
About this role
Pluralis Research is seeking a Research Engineer to tackle the unique challenges of post-training large language models in a decentralized environment. This role involves building an end-to-end RL training stack for consumer-grade devices over the internet, adapting existing algorithms for asynchronous, high-latency scenarios, and ultimately deploying the first post-trained models. The ideal candidate will have hands-on experience with RL post-training, strong engineering skills, and a belief in the potential of Protocol Learning.
What you'll do
- Build the RL training loop, including rollout ingestion, reward computation, and policy updates.
- Adapt existing RL algorithms to handle asynchronous, high-latency, and partially trusted environments.
- Develop evaluation metrics to track model improvement and deploy post-trained models.
- Set the technical direction and drive projects to completion.
- Ship the first decentralized post-trained release.
What they're looking for
- RL Post-training (RLHF, RLVR)
- Python
- PyTorch
- Concurrency
- Failure Handling
- Profiling
- Distributed RL
Benefits
- Equity-Heavy Package
- Remote-First Culture
- Visa Sponsorship (Australia or US)
- Flexible Work Environment
- Open Problems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Pluralis Research
Likely interview questions
- Describe your experience with RL post-training on large language models, including the systems you've worked with.
- How would you approach adapting standard RL algorithms to handle high-latency, asynchronous rollouts?