Tower Research Capital
Machine Learning Research Engineer
About this role
Tower Research Capital seeks an ML Research Engineer to serve as the primary validator and early adopter of its internal machine learning platform. You'll benchmark infrastructure, streamline rapid prototyping workflows, and act as a technical bridge between engineering and research teams, using hands-on testing and AI agents to guide platform development.
What you'll do
- Validate ML infrastructure by running complex models through training and inference pipelines to identify bottlenecks before broader deployment
- Build high-level abstractions and APIs that reduce setup friction for researchers using the ML platform
- Integrate ML tooling with simulation and data frameworks to enable seamless distributed training via Ray
- Design and execute agentic workflows to autonomously generate experiments and stress-test distributed clusters
- Document empirical findings on system performance and hardware capabilities
- Advise engineering and research teams on platform improvements and best practices
What they're looking for
- Python and software design principles
- PyTorch or TensorFlow
- Ray, Dask, or PyTorch Distributed for distributed computing
- System profiling and performance optimization
- LLM tooling and agentic AI frameworks
- Model training, evaluation, and deployment at scale
- API and abstraction design
- Debugging across hardware and software layers
Opens the official application on the employer’s site. No login required.
Tower Research Capital
Tower Research Capital builds high-performance quantitative trading systems and infrastructure, serving traders and researchers with low-latency platforms for market data, strategy execution, and order management. The company is hiring software engineers, quantitative developers, and infrastructure specialists to design scalable systems, optimize trading platforms, and enhance development tooling.
View all jobs at Tower Research CapitalLikely interview questions
- Describe a time you identified and resolved a performance bottleneck in a distributed ML system—what tools did you use and how did you validate the fix?
- How have you approached designing APIs or abstractions that other engineers found intuitive and valuable to use?