Lambda
Data Center Operations Systems Engineer (Seattle)
About this role
Lambda seeks a Data Center Operations Systems Engineer to manage infrastructure deployment, hardware troubleshooting, and inventory at their Quincy, WA facility. You'll rack and cable advanced GPU systems, maintain DCIM documentation, coordinate with supply chain teams, and resolve complex hardware incidents in a fast-growing AI cloud infrastructure company.
What you'll do
- Rack, cable, label, and configure new server, storage, and network infrastructure for GPU and advanced networking systems
- Troubleshoot hardware and software issues in GPU clusters and resolve escalated incidents with hardware support teams
- Maintain data center layout documentation and network topology in DCIM software with accurate tracking
- Manage parts depot inventory and track equipment through delivery, staging, deployment, and handoff processes
- Coordinate with supply chain and manufacturing teams on deployment timelines and large-scale project planning
- Partner with RMA team to process faulty equipment returns and manage replacement orders
What they're looking for
- Data center infrastructure (power distribution, airflow management, environmental monitoring, capacity planning)
- DCIM software and structured cabling systems
- Carrier DIA circuit testing, fiber optics, and cable troubleshooting
- Single and three-phase power theory, PDU balancing, and cold/hot aisle containment
- Server hardware knowledge and boot process understanding
- Cable media types and network topology documentation
- Cross-functional collaboration and MOP development
- Linux administration (nice to have); InfiniBand 400Gb and GPU cluster experience (nice to have)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Describe your experience with DCIM software and how you've used it to maintain data center documentation and layout accuracy.
- Walk us through how you would troubleshoot a GPU system boot failure and determine whether the issue is hardware or firmware-related.