Clera
ML Infrastructure Engineer
About this role
Build and scale inference and model-serving infrastructure for AI agents at an early-stage enterprise AI company. Own the end-to-end production systems ensuring agents run fast and reliably at scale across regulated industries.
What you'll do
- Design and build inference and model-serving infrastructure from ground up through production deployment
- Optimize systems for latency, throughput, and reliability under high concurrency
- Collaborate with ML and infrastructure teams to ensure seamless integration
- Identify and surface performance bottlenecks in production systems
- Drive solutions to infrastructure challenges across cross-functional teams
- Manage ML workloads on cloud platforms and ensure monitoring and observability
What they're looking for
- ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
- Distributed systems and containerization (Docker, Kubernetes)
- Monitoring and observability tools (Prometheus, Grafana, distributed tracing)
- Cloud platform deployment (AWS, GCP, Azure)
- Systems or backend programming (Python, Go, Rust, C++, Java)
- Production ML infrastructure at scale
- Low-latency and real-time systems optimization
- Multi-step AI pipeline orchestration
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a time you designed an inference serving system from scratch—what were the key architectural decisions and trade-offs?
- How have you approached optimizing latency and throughput in a high-concurrency ML serving environment?