Arize AI
DevOps Engineer
About this role
Arize AI, a Series C company backed by $135M in funding, seeks a DevOps Engineer to manage infrastructure for distributed, scalable services across SaaS and on-premises deployments. You'll work hands-on with Kubernetes and multi-cloud environments, collaborating with customers and product teams to ensure reliable, high-performing AI observability platforms.
What you'll do
- Manage and optimize infrastructure supporting distributed services in SaaS and on-prem environments
- Gather customer requirements and adapt deployment manifests for diverse infrastructure needs
- Monitor platform health and performance using observability tools
- Collaborate with product teams to test features and package on-prem releases
- Automate and streamline the release pipeline for faster deployments
- Troubleshoot infrastructure issues and optimize system reliability
What they're looking for
- Kubernetes
- AWS, GCP, and Azure cloud platforms
- Infrastructure as Code
- Release pipeline automation
- Monitoring and observability tools
- On-premises deployment architecture
- Troubleshooting and debugging
- Customer communication and requirements gathering
Benefits
- Medical, dental, and vision coverage
- 401(k) plan
- Unlimited paid time off
- Generous parental leave
- Mental health and wellness support
- WFH stipend for co-working spaces
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Arize AI
Arize AI builds observability and evaluation solutions for production generative AI systems, helping enterprise customers monitor and secure their AI deployments. The company is hiring Forward Deployed AI Engineers to implement these solutions directly with clients, DevSecOps Engineers to secure AI infrastructure, and AI Sales Engineers to guide enterprise customers through technical evaluations and deployments.
- Website
- arize.ai
Likely interview questions
- Describe your experience deploying and managing Kubernetes clusters across multiple cloud providers—how did you handle differences between AWS, GCP, and Azure?
- Tell us about a time you had to troubleshoot a production infrastructure issue; what was your approach and what did you learn?