Lambda
AI Operations Engineer - IT/Internal Infrastructure
San Francisco Office (Fremont St) (Remote)$206k–$275kfulltimemidAdded 5 days ago
About this role
Lambda seeks an IT Systems Engineer to design and operate internal platforms and services supporting the company's AI cloud infrastructure. You'll build scalable, reliable systems across distributed infrastructure while collaborating with multiple teams on architecture, automation, and operational excellence.
What you'll do
- Design and implement software to improve availability, scalability, reliability, and efficiency of internal IT systems
- Solve critical service issues and build automation to prevent recurrence and reduce manual intervention
- Collaborate with engineering teams to influence system designs, architectures, and standards for large-scale systems
- Conduct capacity planning, demand forecasting, performance analysis, and system tuning
- Create comprehensive documentation and architectural artifacts for systems under management
- Participate in on-call rotation and incident response for mission-critical services
What they're looking for
- System design and architecture for performance and scalability
- Cloud infrastructure platforms (AWS, GCP, Azure)
- Configuration management tools (Chef, Ansible, Terraform, GitHub Actions)
- Programming in Python and Go
- Distributed systems thinking and failure mode analysis
- Asynchronous communication and technical documentation
- Incident response and monitoring/alerting systems
- Infrastructure automation and CI/CD practices
Benefits
- Generous cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- 401k plan with 2% company match
- Flexible paid time off
- Wellness and commuter stipends for select roles
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Describe a time you diagnosed and resolved a critical outage in a distributed system. What was your approach?
- How would you design a monitoring and alerting system to catch issues before customers notice them?