Anthropic
Software Engineer, ML Networking
About this role
Join Anthropic as a ML Networking Software Engineer to design and optimize network infrastructure connecting AI accelerators. You'll build low-level networking software, debug distributed systems, and work on high-performance protocols in a fast-growing AI safety organization.
What you'll do
- Design and maintain software interfacing accelerators with high-speed networks
- Develop kernel-space and user-space networking solutions using technologies like DPDK, RDMA, and eBPF
- Diagnose and resolve networking issues in large-scale distributed systems
- Implement and optimize collective algorithms and RPC abstractions for ML workloads
- Benchmark and optimize network performance for synchronous ML training
- Debug kernel-level latency issues and network protocol behavior
What they're looking for
- Network protocols and TCP/IP stack internals
- Kernel networking (XDP, eBPF, io_uring, epoll)
- User-space networking (DPDK, RDMA, kernel bypass)
- Systems programming (C/C++, Rust preferred)
- Performance optimization and profiling
- Distributed systems debugging and design
- ML accelerator architecture
- PCIe and driver development
Benefits
- Work on cutting-edge AI infrastructure at scale
- Hybrid work with 25% office requirement across SF, NYC, or Seattle
- Visa sponsorship support with dedicated immigration attorney
- Mission-driven work on safe, beneficial AI systems
- Competitive annual compensation package
- Collaborative environment with researchers and systems engineers
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anthropic
Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.
- Website
- anthropic.com
Likely interview questions
- Walk us through your experience optimizing network performance in distributed systems—what was the bottleneck and how did you resolve it?
- Describe a time you debugged a complex, multi-layered networking issue across kernel and user space. What was your approach?