Stack AV
Staff Software Engineer, ML Training Infrastructure
About this role
About Stack
Stack is developing revolutionary AI and advanced autonomous systems designed to enhance safety, reliability, and efficiency of modern operations. Stack's autonomous technology incorporates cutting-edge advancements in artificial intelligence, robotics, machine learning, and cloud technologies, empowering us to create innovative solutions that address the needs and challenges of the dynamic trucking transportation industry. With decades of experience creating and deploying real world systems for demanding environments, the Stack team is dedicated to developing an autonomous solution ecosystem tailored to the trucking industry's unique demands.
About the Role
The ML Training team is dedicated to increasing Stack's AV development velocity by accelerating machine learning iterations. Our core mission is to deliver a training system that is reliable, scalable, user-friendly and observable. We are responsible for the overall AI workflow, ranging from dataset curation to training, validation, acceleration, optimization and deployment of large-scale models that power our autonomous vehicles. In addition, this team is in charge of evangelizing best practices and frameworks among Machine Learning Engineers (MLEs) across the company.
In this Staff role, you will drive the design and development of a high-performance multi-tenant AI training platform. You will balance hands-on coding with long-term technical direction by operating across ML Platform, Infrastructure, Autonomy and Safety Evaluation teams to accelerate the development of autonomous vehicles across Stack.
Responsibilities
- Design & evolve high-performance training platform components including orchestration, training abstractions, control plane, observability and performance tuning.
- Deliver end-to-end ML model pipelines across logs processing, feature extraction, dataset schema design/storage, model configuration management, model training, and profiling/acceleration workflows.
- Analyze training infrastructure performance to identify and resolve performance bottlenecks
- Evangelize system abstractions and tooling that enable MLEs to rapidly iterate on models.
- Incorporate OSS tools to enable ML engineers self-sufficiently profile and optimize their workflows
- Promote Engineering Excellence: Maintain a high bar for engineering excellence in their own work but also set a culture of engineering excellence within the team.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 6+ years of experience with ML Platforms and building ML-based applications (modeling experience is a bonus).
- Strong programming skills in Python, C++ or equivalent.
- Prior experience with Lance, PyTorch, Ray Data or equivalent technologies.
- Proven track record of building scalable, reliable infra in a fast-paced environment while working with MLEs across multiple modeling teams.
- A deep understanding of design tradeoffs and ability to articulate those tradeoffs to build alignment across XFN teams.
- Experience with model training, model optimization, or large-scale data processing pipelines.
- Autonomous vehicles (AV) experience is a bonus.
- Strong analytical and problem-solving skills.
- Excellent verbal and written communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.
#LI-AW1
We are proud to be an equal opportunity workplace. We believe that diverse teams produce the best ideas and outcomes. We are committed to building a culture of inclusion, entrepreneurship, and innovation across gender, race, age, sexual orientation, religion, disability, and identity.
Check out our Privacy Policy.
Please Note: Pursuant to its business activities and use of technology, Stack AV complies with all applicable U.S. national security laws, regulations, and administrative requirements, which can restrict Stack AV’s ability to employ certain persons in certain positions pursuant to a range of national security-related requirements. As such, this position may be contingent upon Stack AV verifying a candidate’s residence, U.S. person status, and/or citizenship status. This position may also involve working with software and technologies subject to U.S. export control regulations. Under these regulations, it may be necessary for Stack AV to obtain a U.S. government export license prior to releasing its technologies to certain persons. If Stack AV determines that a candidate’s residence, U.S. person status, and/or citizenship status will require a license, prohibit the candidate from working in this position, or otherwise be subject to national security-related restrictions, Stack AV expressly reserves the right to either consider the candidate for a different position that is not subject to such restrictions, on whatever terms and conditions Stack AV shall establish in its sole discretion, or, in the alternative, decline to move forward with the candidate’s application.
Written by Stack AV. Original job post
Skills mentioned
- C++
- Machine Learning
- Python
- PyTorch
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Stack AV
Stack AV builds infrastructure and systems for autonomous vehicles, focusing on the computational platforms that power their operations at scale. The company is hiring Site Reliability Engineers and infrastructure specialists to optimize performance, reliability, and efficiency across their autonomous systems.
- Website
- stackav.com
Preparing likely interview questions for this role…