Glean
Software Engineer, Compute Infrastructure
About this role
Glean seeks a Software Engineer for Compute Infrastructure to design and operate the Kubernetes-based runtime platform powering AI search, assistant, and agentic workloads. You'll build scalable, cost-efficient multi-cloud systems serving production services at enterprise scale.
What you'll do
- Design and own backend/platform services for runtime infrastructure with focus on reliability, scalability, and AI workload performance
- Develop Kubernetes-based runtime primitives including service orchestration, scheduling, and autoscaling across multi-cloud environments (GCP, AWS, Azure)
- Collaborate with platform, data, and product teams to establish golden paths for service deployment, configuration, and operations
- Drive improvements in latency, resource utilization, and cost for core platform services and multitenant environments
- Implement infrastructure-as-code patterns, observability, and production guardrails including SLOs, dashboards, and safe rollout procedures
What they're looking for
- Kubernetes and container orchestration
- Multi-cloud infrastructure (GCP, AWS, Azure)
- Backend/platform service design
- Infrastructure-as-code
- System performance optimization
- Observability and monitoring
- Production reliability and SLOs
- Distributed systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Glean
Glean builds a Work AI platform that helps enterprises access and leverage their internal data intelligently. The company is hiring backend engineers, infrastructure specialists, fullstack engineers, and billing platform leads to develop scalable features, robust data infrastructure, consumption-based billing systems, and enterprise-grade storage and analytics capabilities.
- Website
- glean.com
Likely interview questions
- Describe your experience designing and operating Kubernetes-based systems in production. What scaling challenges have you solved, and how did you approach them?
- Tell us about a time you optimized infrastructure costs or latency for a multi-cloud or distributed system. What tradeoffs did you make?