Skip to main content

Anyscale

Site Reliability Engineer, Platform Infrastructure (Foundations)

San Francisco (Remote)$200k–$240kfulltimemidAdded today

About this role

Anyscale is hiring a Site Reliability Engineer to build and optimize the infrastructure backbone supporting Ray, an open-source distributed computing platform. You'll design scalable control and data plane systems, manage Kubernetes orchestration, and ensure high-performance execution of AI/ML workloads across cloud and on-premises environments.

What you'll do

  • Design and scale services orchestrating Ray clusters across multiple cloud providers and on-prem environments
  • Optimize control plane components for large-scale distributed AI/ML workloads
  • Build intelligent scheduling and resource management systems for heterogeneous compute clusters
  • Develop features enhancing reliability, performance, scalability, and observability of Ray workloads
  • Support accelerator integration (GPUs, TPUs) and manage container images for distributed systems
  • Provide on-call support and troubleshoot infrastructure issues with customer teams

What they're looking for

  • Kubernetes and container orchestration
  • Cloud platforms (AWS, Azure, GCP)
  • Distributed systems design and implementation
  • High-availability and scalable system architecture
  • Networking, security, and cloud authentication
  • Observability tools (Prometheus, Grafana)
  • Production code development (Python, Go, or similar)
  • Ray framework familiarity
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Anyscale

Anyscale builds Ray, an open-source distributed computing framework and enterprise platform for scaling AI workloads across Kubernetes and cloud providers. The company is hiring forward-deployed engineers to work embedded with customers, software engineers to develop Ray Core, LLM inference specialists, and customer support engineers who combine technical expertise with post-sale success.

View all jobs at Anyscale

Likely interview questions

  • Describe your experience designing and maintaining highly available distributed systems in production.
  • How have you optimized control plane performance when handling large-scale workloads?