Skip to main content

Sciforium

GPU Cluster Engineer, Networking

  • Confirmed live in the last 24 hours
  • $150k–$180k
  • Mid level
  • Full-time
  • On-site · San Francisco, CA
  • 7+ yrs exp
  • Added 2 months ago

About this role

Sciforium is seeking a Senior Network Engineer to design, build, and operate the networking infrastructure for their rapidly expanding GPU clusters. This role will own the entire network stack, from the data center perimeter to cross-site connectivity, ensuring high performance and reliability for AI workloads. The ideal candidate will possess deep expertise in RDMA, network automation, and cloud networking.

What you'll do

  • Architect full network designs for new GPU clusters, including topology and growth planning.
  • Configure and optimize RDMA fabrics (InfiniBand/RoCE v2) for training and inference workloads.
  • Implement network security measures including firewalls, VPNs, and network segmentation.
  • Automate network configuration and deployment using tools like Ansible and Git.
  • Troubleshoot and resolve network performance issues and incidents.
  • Design and manage cross-cluster and hybrid cloud connectivity solutions.

What they're looking for

  • RDMA (InfiniBand, RoCE v2)
  • BGP, OSPF, EVPN-VXLAN
  • Ansible, Nornir, NAPALM
  • Network Security (Firewalls, VPNs)
  • Python
  • Network Automation
  • Cloud Networking (AWS, Azure, GCP)

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary
  • Equity
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Sciforium

View all jobs at Sciforium