Lambda
Data Center Operations Systems Engineer III (Los Angeles)
About this role
Lambda seeks a Data Center Operations Systems Engineer III to manage GPU and advanced networking infrastructure deployment, maintenance, and troubleshooting at their Los Angeles data center. This role involves hands-on hardware operations, inventory management, and collaboration across teams to support large-scale AI cloud infrastructure, with 5 days/week on-site presence and shift work.
What you'll do
- Rack, cable, label, and configure new servers, storage, and network infrastructure
- Troubleshoot hardware and software issues in GPU and advanced networking systems
- Update data center layout and network topology in DCIM software
- Manage parts depot inventory and track equipment through deployment lifecycle
- Collaborate with hardware support teams on complex troubleshooting and incident resolution
- Partner with RMA team on faulty equipment returns and replacement ordering
What they're looking for
- Data center infrastructure (power distribution, airflow, environmental monitoring, DCIM)
- Structured cabling and cable management practices
- Fiber optics, DIA circuit testing, and network troubleshooting
- Single and three-phase power theory and PDU balancing
- Server hardware knowledge and boot processes
- Cold aisle and hot aisle containment principles
- Process documentation and cross-team collaboration
- Linux administration (preferred)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Describe your hands-on experience with data center infrastructure—specifically power distribution, thermal management, and DCIM tools you've used.
- Walk us through how you've diagnosed and resolved a complex networking or hardware issue in a production data center environment.