Lambda
Data Center Operations Systems Engineer (Atlanta)
About this role
Lambda seeks a Data Center Operations Systems Engineer to manage infrastructure deployment, maintenance, and troubleshooting at their Atlanta facility supporting AI cloud services. The role involves racking and configuring servers, managing inventory, coordinating with supply chain teams, and ensuring compliance with installation standards across large-scale deployments.
What you'll do
- Rack, cable, label, and configure new server, storage, and network infrastructure
- Troubleshoot hardware and software issues in advanced data center systems
- Document data center layout and network topology using DCIM software
- Manage parts depot inventory and track equipment through deployment lifecycle
- Coordinate with supply chain and manufacturing teams on large-scale deployment projects
- Collaborate with hardware support and RMA teams to resolve infrastructure issues
What they're looking for
- Data center critical infrastructure (power distribution, cooling, environmental monitoring)
- DCIM software proficiency
- Structured cabling and cable management
- Server hardware troubleshooting
- Network topology understanding
- Linux administration
- Ticketing systems (JIRA, Zendesk)
- Attention to detail and documentation
Benefits
- Health, dental, and vision coverage for you and dependents
- Equity and cash compensation
- 401k plan with 2% company match
- Flexible paid time off
- Wellness and commuter stipends
- Opportunity to work with cutting-edge AI infrastructure
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lambda
Lambda builds AI cloud infrastructure providing GPU compute and networking capabilities for researchers and enterprises. The company is hiring for data center operations, security, and facility engineering roles to support large-scale AI compute deployments.
View all jobs at LambdaLikely interview questions
- Describe your experience with racking and cabling infrastructure at scale—how do you ensure consistency and proper documentation?
- Tell us about a time you troubleshot a complex hardware issue in a data center environment. How did you approach it?