Tower Research Capital
GPU Systems Engineer
About this role
Tower Research Capital seeks a GPU Systems Engineer to design, deploy, and operate large-scale distributed GPU clusters supporting global trading and AI research. The role spans infrastructure architecture, performance optimization, and automation across thousands of nodes handling petabyte-scale workloads.
What you'll do
- Design and deploy distributed GPU clusters including hardware selection, network topology, and production operations at scale
- Identify and resolve performance bottlenecks across compute, storage, network, and interconnect layers
- Profile GPU workloads with researchers and optimize for measurable performance improvements
- Build automation for cluster provisioning, monitoring, diagnostics, and self-healing across thousands of nodes
- Own infrastructure projects end-to-end from design through implementation and ongoing support
- Evaluate new hardware and software generations, collaborating with vendors on complex issues
What they're looking for
- Large-scale Linux systems administration in HPC or distributed infrastructure environments
- Linux kernel tuning, performance debugging, and system-level troubleshooting
- GPU performance modeling and distributed GPU workload troubleshooting
- GPUDirect RDMA and GPU-network data movement optimization
- Python for automation and scripting, plus CUDA or C/C++ proficiency
- Configuration management tools (Ansible, Salt, Puppet, or Chef)
- Cross-layer diagnostics spanning hardware, OS, and networking
- NVIDIA stack familiarity (NCCL, NVLink)
Benefits
- Competitive base salary $200,000–$300,000 plus discretionary bonus
- Generous paid time off policies
- Hybrid working arrangements
- Free daily meals and well-stocked kitchens
- Wellness reimbursement and in-office fitness programs
- Company-sponsored sports teams and volunteer opportunities
Opens the official application on the employer’s site. No login required.
Tower Research Capital
Tower Research Capital builds high-performance quantitative trading systems and infrastructure, serving traders and researchers with low-latency platforms for market data, strategy execution, and order management. The company is hiring software engineers, quantitative developers, and infrastructure specialists to design scalable systems, optimize trading platforms, and enhance development tooling.
View all jobs at Tower Research CapitalLikely interview questions
- Describe a time you debugged a performance issue that spanned multiple layers—hardware, OS, and network. How did you isolate the root cause?
- Walk us through your approach to profiling a GPU workload that isn't scaling as expected. What tools and metrics do you rely on?