Clera
ML Infrastructure Engineer
About this role
Seed-stage AI infrastructure company seeks an ML Infrastructure Engineer to design and operate inference and model-serving systems that power reliable AI agents for regulated industries. You'll own production ML infrastructure end-to-end, optimizing for latency, throughput, and reliability as the platform scales.
What you'll do
- Design and deploy inference and model-serving infrastructure from initial architecture through production scaling
- Build systems enabling AI agents to run reliably under high concurrency and increasing demand
- Identify and resolve infrastructure bottlenecks in collaboration with ML and platform teams
- Optimize production systems for latency, throughput, and reliability in cloud environments
- Establish observability, monitoring, and debugging practices across the ML stack
What they're looking for
- ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
- Containerization and orchestration (Docker, Kubernetes)
- Distributed systems design with high concurrency handling
- Monitoring and observability tools (Prometheus, Grafana, ELK, distributed tracing)
- Cloud platform deployment (AWS, GCP, Azure)
- Systems or backend languages (Python, Go, Rust, C++, Java)
- Production ML system optimization
- Real-time inference and low-latency serving
Benefits
- Early-stage ownership and significant impact on foundational infrastructure
- Work on complex distributed systems problems in production AI
- Small, experienced team with deep ML and enterprise engineering expertise
- Well-funded seed-stage company with strong institutional backing
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a time you optimized an ML inference system for latency under high concurrency—what was the bottleneck and how did you resolve it?
- Describe your experience with model-serving platforms like TorchServe, Triton, or KServe and which you'd choose for a multi-tenant production system and why.