Tower Research Capital
Machine Learning Performance Engineer, Training
About this role
Bridge quantitative research and high-performance computing by optimizing machine learning model training at scale. You will accelerate the full training lifecycle—from data ingestion through kernel execution—enabling researchers to iterate faster on complex models and datasets across CPU, GPU, and specialized accelerator platforms.
What you'll do
- Benchmark model-training workloads across CPUs, GPUs, and accelerators to identify bottlenecks and guide infrastructure decisions
- Design and optimize distributed training strategies including data, tensor, pipeline, and model parallelism
- Develop and optimize GPU kernels and performance-critical framework components for ML workloads
- Analyze and improve the full training pipeline from data loading through checkpointing to increase accelerator utilization
- Apply numerical optimization techniques such as mixed-precision training, activation checkpointing, and operator fusion
- Partner with HPC and infrastructure teams to optimize workload scheduling, resource allocation, and fault tolerance
What they're looking for
- PyTorch or JAX framework optimization and execution models
- GPU kernel development (CUDA, Triton, CUTLASS, cuBLAS, cuDNN)
- Python and C++ programming
- Distributed-training technologies (NCCL, FSDP, DeepSpeed, Megatron-LM)
- Performance analysis tools (Nsight Systems/Compute, PyTorch Profiler)
- GPU architecture and memory hierarchy knowledge
- High-performance networking and interconnects (InfiniBand, RDMA, NVLink)
- Benchmarking and data-driven performance analysis
Opens the official application on the employer’s site. No login required.
Tower Research Capital
Tower Research Capital builds high-performance quantitative trading systems and infrastructure, serving traders and researchers with low-latency platforms for market data, strategy execution, and order management. The company is hiring software engineers, quantitative developers, and infrastructure specialists to design scalable systems, optimize trading platforms, and enhance development tooling.
View all jobs at Tower Research CapitalLikely interview questions
- Walk us through how you've identified and resolved a performance bottleneck in a distributed training system—what tools did you use and what was the outcome?
- Describe your experience optimizing GPU kernels. Which libraries have you worked with, and how did you balance performance gains against development complexity?