Xometry
Site Reliability Engineer II (SRE)
About this role
Xometry seeks a Site Reliability Engineer II to drive infrastructure reliability and performance across engineering teams. You'll own technical problem statements end-to-end, design scalable platform solutions, and guide deployment practices while balancing safety with speed.
What you'll do
- Own assigned infrastructure and reliability problem statements from conception through completion
- Develop and maintain cloud platforms, Kubernetes clusters, and AWS infrastructure for production systems
- Build and configure observability and monitoring solutions using tools like Coralogix and Sentry
- Develop CI/CD infrastructure and tooling using GitHub Actions and ArgoCD
- Write clean, well-documented code while improving existing systems and learning from feedback
- Collaborate across teams with clear communication on progress, blockers, and technical outcomes
What they're looking for
- Python, JavaScript, or Unix Shell scripting
- AWS infrastructure and production workload management
- Kubernetes cluster configuration and maintenance
- CI/CD pipeline development and GitHub Actions
- Observability and monitoring tool implementation
- Infrastructure-as-code and cloud architecture
- Problem-solving and technical project ownership
- Cross-team collaboration and communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Xometry
Xometry operates a manufacturing network platform that connects with suppliers and partners to deliver complex projects, particularly in regulated industries like aerospace and defense. The company is hiring analytics engineers to build data infrastructure, and supplier quality and development engineers to optimize partner performance, ensure product quality, and manage manufacturing relationships.
- Website
- xometry.com
Likely interview questions
- Describe your experience managing and scaling production workloads on AWS—what challenges did you face?
- How have you approached designing a Kubernetes infrastructure to balance reliability with operational simplicity?