Sciforium
GPU Cluster Engineer, Networking
- Confirmed live in the last 24 hours
- $150k–$180k
- Mid level
- Full-time
- On-site · San Francisco, CA
- 7+ yrs exp
- Added 2 months ago
About this role
Sciforium is seeking a Senior Network Engineer to design, build, and operate the networking infrastructure for their rapidly expanding GPU clusters. This role will own the entire network stack, from the data center perimeter to cross-site connectivity, ensuring high performance and reliability for AI workloads. The ideal candidate will possess deep expertise in RDMA, network automation, and cloud networking.
What you'll do
- Architect full network designs for new GPU clusters, including topology and growth planning.
- Configure and optimize RDMA fabrics (InfiniBand/RoCE v2) for training and inference workloads.
- Implement network security measures including firewalls, VPNs, and network segmentation.
- Automate network configuration and deployment using tools like Ansible and Git.
- Troubleshoot and resolve network performance issues and incidents.
- Design and manage cross-cluster and hybrid cloud connectivity solutions.
What they're looking for
- RDMA (InfiniBand, RoCE v2)
- BGP, OSPF, EVPN-VXLAN
- Ansible, Nornir, NAPALM
- Network Security (Firewalls, VPNs)
- Python
- Network Automation
- Cloud Networking (AWS, Azure, GCP)
Benefits
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary
- Equity
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.