Ooma
Site Reliability Engineer
Remote, US (Remote)From $170kmidAdded today
About this role
Ooma seeks an experienced Site Reliability Engineer to manage and optimize large-scale on-premises data center infrastructure across multiple US locations. You'll oversee bare metal servers, virtualization platforms, Kubernetes clusters, and network infrastructure while maintaining system reliability and implementing automation at scale.
What you'll do
- Manage hundreds of bare metal servers and VMs across data centers, handling hardware lifecycle, firmware updates, and vendor coordination
- Monitor system performance and troubleshoot issues using observability tools, with focus on OS and bare metal diagnostics
- Automate bare metal and VM provisioning, OS lifecycle management, and configuration using Ansible across the fleet
- Design and operate Kubernetes clusters on-premises and manage high-throughput Kafka clusters for event streaming
- Administer network infrastructure including VLANs, routing, load balancers, and firewalls; troubleshoot latency and connectivity
- Participate in on-call rotations, conduct post-mortems, and maintain runbooks and technical documentation
What they're looking for
- Linux systems administration and troubleshooting
- Bare metal server management and hardware lifecycle
- Virtualization platforms and storage administration
- Kubernetes and container orchestration
- Ansible for configuration management and automation
- Kafka cluster design and operation
- Data center networking (VLANs, routing, load balancing)
- DNS, DHCP, NTP, and authentication services
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Ooma
Ooma builds cloud-based communication products and is hiring for quality assurance and engineering roles focused on testing, reliability, and product performance.
- Website
- ooma.com
Likely interview questions
- Describe your experience managing large-scale bare metal server environments and how you've handled hardware failures and RMA processes.
- How have you designed and maintained Kubernetes clusters in on-premises or hybrid environments?