Lambda
Software Engineer - Compute
San Francisco Office (Fremont St) (Remote)$266k–$395kfulltimemidAdded today
About this role
Lambda seeks a Software Engineer to design and maintain GPU-first cloud infrastructure for bare-metal and VM provisioning, focusing on compute instance lifecycle management and system reliability. You'll develop distributed systems for orchestrating compute resources and collaborate across teams to scale engineering productivity.
What you'll do
- Design and develop software for GPU/CPU compute infrastructure with emphasis on performance, scalability, and reliability
- Implement services for bare-metal and virtual machine instance provisioning and lifecycle management
- Build distributed systems for managing and orchestrating compute resources across multiple hardware SKUs
- Troubleshoot and debug complex production and development environment issues
- Participate in on-call rotation and own incident response
- Drive technical discussions and collaborate across multiple teams on architecture and solutions
What they're looking for
- Go (Golang) or Python in production environments
- Bare metal and virtualization hardware management
- Linux system administration and OS-level debugging
- Networking and hardware troubleshooting
- Distributed systems design
- GPU infrastructure or high-performance computing (nice to have)
- Kubernetes or Slurm cluster management (nice to have)
- Virtualization technologies like KVM/QEMU (nice to have)
Benefits
- Generous cash and equity compensation
- Health, dental, and vision coverage for you and dependents
- 401k plan with 2% company match
- Wellness and commuter stipends
- Flexible paid time off
- 4-day office presence requirement with one designated work-from-home day
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Walk us through a complex infrastructure issue you debugged at the OS, hardware, or networking layer—what was your approach?
- Describe your experience building or maintaining bare-metal provisioning systems. What were the key challenges?