Skip to main content

Lambda

Software Engineer - Fleet Orchestration

San Francisco Office (Fremont St) (Remote)$266k–$395kfulltimemidAdded today

About this role

Lambda seeks a Software Engineer to build fleet orchestration systems that manage GPU datacenter operations at scale. You'll design automation for cluster deployment, state validation, and cross-team coordination while owning production SLAs in a critical infrastructure role.

What you'll do

  • Build fleet data systems to index, validate, and reconcile physical GPU host state against intended logical configuration
  • Design and implement automation for GPU cluster deployments from logical design through OS provisioning and validation
  • Create systems to continuously validate consistency between intended and actual state, detecting drift before failures
  • Implement ownership, global locking, readiness gating, and safety systems for coordinated fleet operations
  • Lead technical contribution to new datacenter site bring-ups across multiple teams
  • Monitor production health, maintain SLAs, and drive resolution of fleet drift issues

What they're looking for

  • Python or Go programming
  • Distributed systems design
  • PXE boot, firmware provisioning, and IPMI/BMC interfaces
  • Network services (DNS, DHCP)
  • Technical design and documentation
  • Cross-functional collaboration and influence
  • Production systems ownership and mentoring
  • Infrastructure automation and APIs

Benefits

  • Generous cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends
  • 401k plan with 2% company match
  • Flexible paid time off
  • Work from home one day per week (Tuesdays)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Lambda

Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.

View all jobs at Lambda

Likely interview questions

  • Describe a time you owned a production system with real SLAs—how did you monitor it and respond to incidents?
  • Walk us through how you'd approach designing a fleet validation system that detects drift between intended and actual state across thousands of hosts.