Lambda
Technical Success Engineer
About this role
Lambda seeks a Technical Success Engineer to shepherd enterprise customers from signed deployment through production readiness on their AI cloud infrastructure. You'll validate complex GPU and Kubernetes deployments, coordinate across engineering teams to resolve blockers, and serve as the technical lead through customer onboarding and early production operations.
What you'll do
- Manage deployments from contract signature to live production, validating compute, storage, connectivity, and configuration against commitments
- Troubleshoot and resolve technical gaps by coordinating with engineering, infrastructure, product, and data center teams
- Maintain accurate deployment status tracking and deliver regular stakeholder updates with RAG ratings and risk assessments
- Guide customers through onboarding and first successful production workload execution
- Serve as primary technical contact and escalation point during deployment and early production phases
- Document recurring technical patterns into reusable runbooks and automation for future engagements
What they're looking for
- GPU/HPC infrastructure and large-scale Linux systems
- Kubernetes and cloud platform administration
- Network, storage, and compute troubleshooting
- Technical project coordination and cross-team collaboration
- Clear written and verbal communication
- Configuration validation and gap analysis
- Structured status reporting and project tracking
- Hands-on deployment and production operations
Benefits
- Generous cash and equity compensation
- Work in fast-growing AI infrastructure company backed by NVIDIA, ARK Invest, and top-tier investors
- 4 days per week on-site in San Francisco office; Tuesday is designated work-from-home day
- Opportunity to work with enterprise and hyperscaler customers at scale
- Technical leadership role with ownership and autonomy
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Walk us through a complex infrastructure deployment you've owned end-to-end—what were the biggest technical gaps you uncovered and how did you drive resolution?
- Describe your experience troubleshooting GPU cluster issues. What was the hardest problem you've debugged and how did you approach it?