Discord
Software Engineer, Distributed Systems
About this role
Join Discord's Realtime Infrastructure team to design and operate mission-critical distributed systems that power real-time communication for millions of users. This role combines backend development with infrastructure management to ensure platform reliability and enable new features at scale.
What you'll do
- Build and maintain large-scale, reliable distributed systems for text chat and session updates
- Collaborate with product teams to design and implement new platform features
- Operate and manage tier 0 critical services ensuring high availability
- Implement monitoring, alerting, and observability solutions for complex systems
- Debug and resolve infrastructure issues across production environments
- Contribute to code reviews and architectural decisions with the team
What they're looking for
- Backend systems design and development
- Distributed systems architecture
- Production infrastructure operations
- Monitoring and alerting implementation
- Open source software familiarity
- Cloud platforms (GCP, AWS)
- Elixir or Rust programming
- DevOps tools (Terraform, Kubernetes, Salt)
Benefits
- Equity compensation
- Comprehensive benefits package
- Relocation assistance available
- Work with talented engineering team
- Direct impact on platform serving millions of users
- San Francisco Bay Area location
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Discord
Discord builds a real-time communication platform serving millions of daily active users, with infrastructure handling trillions of messages across gaming and community use cases. The company is hiring backend engineers, infrastructure specialists, and data engineers to develop distributed systems, database infrastructure, developer tools, and large-scale data platforms that power its core platform.
- Website
- discord.com
Likely interview questions
- Tell us about a complex distributed system problem you solved and your approach to debugging it.
- How have you ensured high availability and reliability in a tier 0 service you've operated?