Havoc AI
ML Cloud Infrastructure Engineer
About this role
Build and operate cloud infrastructure for machine learning at Havoc, developing pipelines, platforms, and services that connect data, training, evaluation, and deployment across AWS and Kubernetes. You'll enable autonomy and ML teams to move efficiently from raw data to production models while maintaining scalability, reliability, and security.
What you'll do
- Design and operate scalable AWS infrastructure and Kubernetes/EKS workloads using Infrastructure as Code
- Build data pipelines that transform multi-modal telemetry, imagery, and sensor data into versioned training datasets
- Develop reproducible training and evaluation workflows with experiment tracking, model versioning, and deployment infrastructure
- Create self-service tooling and monitoring for compute scheduling, training, inference, and cost optimization
- Establish quality signals, regression testing frameworks, and observability across data pipelines and deployed models
- Partner with autonomy, software, data, and security teams to integrate edge capture through cloud training and deployment
What they're looking for
- AWS infrastructure and Infrastructure as Code (Terraform, CloudFormation)
- Kubernetes and container orchestration (EKS, Docker)
- ML pipelines and workflow tools (MLflow, Kubeflow, or similar)
- Data engineering and pipeline development
- Python or similar backend languages
- CI/CD and release automation
- Observability, monitoring, and logging systems
- Cloud security and IAM practices
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Havoc AI
Havoc AI builds autonomous systems software and hardware for mission-critical operations across military and commercial applications in sea, air, and land domains. The company is hiring frontend engineers, embedded software engineers, mission software engineers, mechanical engineers, and business analytics engineers to develop real-time control systems, operational interfaces, and autonomous capabilities for uncrewed vehicles.
View all jobs at Havoc AILikely interview questions
- Describe your experience building ML infrastructure and data pipelines—what tools and frameworks have you used?
- Walk us through how you'd design a scalable training pipeline for multi-modal data (imagery, telemetry, video) on AWS and Kubernetes.