Anyscale
Site Reliability Engineer, Platform Infrastructure (Foundations)
About this role
Anyscale is hiring a Site Reliability Engineer to build and optimize the infrastructure backbone supporting Ray, an open-source distributed computing platform. You'll design scalable control and data plane systems, manage Kubernetes orchestration, and ensure high-performance execution of AI/ML workloads across cloud and on-premises environments.
What you'll do
- Design and scale services orchestrating Ray clusters across multiple cloud providers and on-prem environments
- Optimize control plane components for large-scale distributed AI/ML workloads
- Build intelligent scheduling and resource management systems for heterogeneous compute clusters
- Develop features enhancing reliability, performance, scalability, and observability of Ray workloads
- Support accelerator integration (GPUs, TPUs) and manage container images for distributed systems
- Provide on-call support and troubleshoot infrastructure issues with customer teams
What they're looking for
- Kubernetes and container orchestration
- Cloud platforms (AWS, Azure, GCP)
- Distributed systems design and implementation
- High-availability and scalable system architecture
- Networking, security, and cloud authentication
- Observability tools (Prometheus, Grafana)
- Production code development (Python, Go, or similar)
- Ray framework familiarity
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anyscale
Anyscale builds Ray, an open-source distributed computing framework and enterprise platform for scaling AI workloads across Kubernetes and cloud providers. The company is hiring forward-deployed engineers to work embedded with customers, software engineers to develop Ray Core, LLM inference specialists, and customer support engineers who combine technical expertise with post-sale success.
- Website
- anyscale.com
Likely interview questions
- Describe your experience designing and maintaining highly available distributed systems in production.
- How have you optimized control plane performance when handling large-scale workloads?