Skip to main content

FluidStack

Network Engineer, Design & Engineering

New York, NY (Remote)$202k–$261kfulltimemidAdded 1 month ago

About this role

Design and engineer large-scale network infrastructure for AI training clusters at Fluidstack, a company building civilization-scale compute capacity. You'll own the complete network design lifecycle from customer requirements through deployment, creating lossless Ethernet fabrics and architectures for 100k+ accelerator clusters.

What you'll do

  • Design end-to-end network architectures for AI training and inference workloads across multiple GPU platforms
  • Create topology designs, IP addressing schemes, routing policies, and fabric configurations for front-end, back-end, and storage networks
  • Develop lossless RDMA (RoCEv2) fabrics with PFC, ECN, and congestion management tuning
  • Translate logical designs into physical reality including rack layouts, cabling, power constraints, and airflow considerations
  • Produce high-level and low-level design documents, BOMs, cabling matrices, and lead design reviews
  • Collaborate across hardware, operations, software, and validation teams to ensure designs are buildable and operationally sound

What they're looking for

  • Data center network fabric design and deployment experience
  • Deep expertise in CLOS/fat-tree topologies, BGP, EVPN/VXLAN
  • Lossless RDMA and RoCEv2 fabric design with congestion management
  • First-principles problem-solving and novel architecture reasoning
  • L1-L3 networking fundamentals and tradeoff analysis
  • Multi-vendor GPU platform integration
  • Large-scale cluster architecture (100k+ accelerators preferred)
  • Source-of-truth-driven design generation and automation

Benefits

  • Competitive salary and equity compensation
  • Retirement or pension plan aligned with local standards
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

FluidStack

FluidStack builds AI infrastructure at scale, developing data centers and warehouse operations designed to handle gigawatt-capacity compute deployment. The company is hiring for warehouse engineers, data center operations specialists, product engineers, and people leaders to support rapid infrastructure expansion across multiple sites.

View all jobs at FluidStack

Likely interview questions

  • Walk us through a data center network fabric you designed end-to-end, from requirements to deployment. What were the key tradeoffs you made and why?
  • How would you design a lossless RDMA fabric for a 100k+ accelerator GPU cluster, and what congestion management strategies would you implement?