Instacart
Site Reliability Engineer II
About this role
Join Instacart's Site Reliability Engineering team as an SRE II to ensure platform reliability and performance across large-scale distributed systems. You'll monitor infrastructure, respond to incidents, develop automation, and collaborate with senior engineers to maintain high-availability services in a supportive, growth-focused environment.
What you'll do
- Monitor systems and respond to alerts, escalating issues as needed
- Participate in incident management following established protocols and documentation procedures
- Develop and maintain automation scripts and tools to reduce manual operational tasks
- Support application deployments and ensure reliable service releases
- Maintain and improve documentation for processes, procedures, and systems
- Troubleshoot technical issues in collaboration with senior engineers
What they're looking for
- Software engineering fundamentals (2-4 years experience)
- Programming and scripting languages
- System administration and troubleshooting
- Problem-solving and analytical thinking
- Incident management and response
- Ruby or Go (preferred)
- Cloud platforms (AWS, GCP, Azure preferred)
- Team collaboration and communication
Opens the official application on the employer’s site. No login required.
Instacart
Instacart builds a grocery delivery marketplace and enterprise AI solutions that optimize shopping experiences through machine learning, pricing, recommendations, and generative AI. The company is hiring PhD-level machine learning researchers, systems engineers, and forward-deployed engineers to develop scalable AI systems and deploy agentic solutions for retail and CPG partners.
- Website
- instacart.com
Likely interview questions
- Walk us through your experience responding to and managing production incidents—what was your role and what did you learn?
- Describe a time you wrote an automation script or tool that significantly reduced manual work. What problem did it solve?