DRW
Service Reliability Engineer
About this role
DRW's Container Infrastructure team is hiring a Service Reliability Engineer to manage and optimize their large-scale Kubernetes platform serving trading desks and enterprise teams. You'll focus on platform reliability, cost optimization, automation of manual tasks, and cross-team collaboration to keep hundreds of clusters and thousands of nodes running smoothly.
What you'll do
- Administer and maintain hundreds of Kubernetes clusters across on-premises and cloud environments
- Partner with business units on application onboarding and deployment to the platform
- Automate repetitive operational tasks to improve efficiency and reduce manual work
- Monitor platform performance, costs, and compute consumption per desk
- Develop new platform components and features to meet evolving business needs
- Provide customer service and technical support to internal platform users
What they're looking for
- Kubernetes cluster administration and troubleshooting
- Networking, storage, and containerization fundamentals
- AWS and/or Google Cloud Platform
- Infrastructure automation and scripting
- On-premise infrastructure management
- Written and verbal communication across technical teams
- Problem-solving under constraints
- Kubernetes operators (Flux, Prometheus, Fluentd, Fluentbit)
Benefits
- Group medical, pharmacy, dental, and vision insurance
- 401k with discretionary employer match
- Short and long-term disability coverage
- Life and AD&D insurance
- Health savings accounts and flexible spending accounts
- Annual discretionary bonus eligibility
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
DRW
DRW is a diversified trading firm that builds and maintains trading technology infrastructure, risk management systems, and systematic trading platforms supporting global 24/7 operations across multiple asset classes. The company is hiring Desktop Systems Engineers, Trade Systems Engineers, and Software Engineers to support endpoint infrastructure, trading system reliability, risk analytics, and full-stack trading platform development.
- Website
- drw.com
Likely interview questions
- Tell us about your experience administering Kubernetes clusters at scale—what was the largest environment you've managed?
- Describe a time you automated a repetitive manual task; what tools did you use and what was the impact?