stripe
Software Engineer, High Availability and Disaster Recovery
Seattle, WA$206.1k–$285.6kfull-timemidAdded today
About this role
Stripe seeks a Software Engineer to design and build highly available, globally distributed systems that handle latency-critical workloads and ensure data resilience across multiple regions and data centers.
What you'll do
- Develop global architecture combining lower-availability components into a resilient, highly available system
- Build latency-critical and data redundancy solutions for distributed production environments
- Investigate and resolve production issues in live systems, including database concerns and cross-region optimization
- Design and implement disaster detection and automated failover systems
- Write and deploy solutions using Ruby, Java, Mongo, Postgres, and related technologies
- Collaborate across teams to establish multi-region architecture patterns and high durability best practices
What they're looking for
- Go, Ruby, or Java programming
- Database internals (data movement, leader election, quorums)
- AWS cloud services (S3, EBS, EC2, VPCs)
- Temporal workflow design and implementation
- Trino SQL for data analysis
- Fault-tolerant system design with sharding and replication
- Distributed systems debugging
- High availability architecture
Benefits
- Equity compensation
- Company bonus or sales commissions
- 401(k) plan
- Medical, dental, and vision benefits
- Wellness stipends
- 50% telecommuting flexibility
Opens the official application on the employer’s site. No login required.
stripe
Stripe builds payment infrastructure and financial services platforms, offering APIs and tools that enable developers and businesses to process transactions, detect fraud, verify identity, and manage security at scale. The company is hiring Backend Engineers, Full Stack Engineers, ML Engineers, AI Engineers, and Security Engineers to develop core platform systems, payment intelligence, customer support infrastructure, and security data platforms.
- Website
- stripe.com
Likely interview questions
- Walk us through your experience designing fault-tolerant systems using data sharding and replication strategies.
- Tell us about a time you investigated and resolved a critical issue in a live, distributed production system.