Skip to main content

Rubrik Job Board

Software Engineer - Reliability (US Citizen)

Palo Alto, CAFrom $237kmidAdded today

About this role

Rubrik is seeking a Site Reliability Engineer to ensure high availability and performance of their Polaris Cloud Platform infrastructure, with a focus on database systems and FedRAMP compliance. You'll manage Kubernetes and MySQL environments, drive reliability improvements, and participate in on-call rotations supporting enterprise customers.

What you'll do

  • Maintain high availability and durability of production databases while minimizing customer downtime
  • Design, implement, and optimize relational database systems for performance and reliability at scale
  • Manage backend infrastructure including Kubernetes, MySQL, and related systems
  • Drive FedRAMP certification processes and compliance initiatives
  • Participate in follow-the-sun on-call rotations and debug production issues
  • Build monitoring tools and automation to improve team efficiency

What they're looking for

  • Relational database design and architecture (MySQL, high-availability, disaster recovery)
  • Kubernetes and container orchestration
  • Programming in Golang, Python, Java, Scala, or C++
  • Distributed systems design and troubleshooting
  • Google Cloud Platform or public cloud experience
  • Linux/Unix systems and networking
  • FedRAMP certification knowledge
  • SQL query optimization and performance tuning
Apply on the employer's site

Opens the official application on the employer’s site. No login required.

Rubrik Job Board

Rubrik builds Disaster Recovery as a Service solutions across cloud platforms with a focus on innovative backend systems and security. The company is hiring Software Engineers, Sales Engineers, and Application Security Engineers to expand its engineering, sales, and security capabilities.

View all jobs at Rubrik Job Board

Likely interview questions

  • Describe your experience designing and operating database systems for large-scale SaaS products with high SLO/SLA requirements.
  • How have you approached database upgrades or migrations while minimizing downtime for production customers?