Skip to main content

Pluralis Research

Machine Learning Engineer - ML Training Platform

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • Remote · San Francisco
  • Added 1 month ago

About this role

Pluralis Research is seeking a Machine Learning Engineer to build and scale a platform for decentralized model training on consumer-grade devices. This role focuses on architecting robust, multi-cloud infrastructure and distributed systems capable of handling real-world network conditions. The ideal candidate will be passionate about Protocol Learning and thrive in a fast-paced, remote-first startup environment.

What you'll do

  • Design resource management systems across AWS, GCP, and Azure.
  • Architect fault-tolerant infrastructure for distributed ML training.
  • Build systems to simulate and handle real network conditions (bandwidth shaping, latency).
  • Manage node churn and ensure data flow across heterogeneous networks.
  • Implement infrastructure-as-code using Pulumi or Terraform.
  • Develop robust retry strategies and health monitoring.

What they're looking for

  • Infrastructure-as-Code (Pulumi/Terraform)
  • Docker/Kubernetes (EKS)
  • GPU Workloads
  • Distributed Systems
  • Python (asyncio, concurrency)
  • Prometheus/Grafana
  • Networking (P2P, NAT traversal)
  • Cloud SDKs

Benefits

  • Equity-Heavy Package
  • Remote-First Culture
  • Visa Sponsorship (Australia or US)
  • Relocation Support (Australia or US)
  • Flexible work environment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Pluralis Research

View all jobs at Pluralis Research

Likely interview questions

  • Describe your experience with infrastructure-as-code tools like Pulumi or Terraform, and a specific challenge you overcame.
  • Explain your understanding of distributed training workflows, including checkpointing and data sharding.