Clera
ML Infrastructure Engineer
About this role
Own the end-to-end inference and model-serving infrastructure for AI agents at an early-stage enterprise AI company. Design scalable, reliable systems that support production workloads in regulated industries while optimizing performance under growing concurrent load.
What you'll do
- Design and build inference and model-serving infrastructure from architecture through production deployment
- Scale systems to handle increasing concurrent load while maintaining reliability and efficiency
- Identify and resolve infrastructure bottlenecks in collaboration with ML and platform teams
- Drive performance optimization across latency, throughput, and reliability metrics
- Manage containerization and orchestration of ML workloads using Docker and Kubernetes
- Implement monitoring and observability solutions for production systems
What they're looking for
- ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
- Distributed systems design and concurrent request management
- Docker and Kubernetes orchestration
- Cloud platforms (AWS, GCP, or Azure)
- Monitoring and observability tools (Prometheus, Grafana, ELK, distributed tracing)
- Systems or backend programming (Python, Go, Rust, C++, or Java)
- Knowledge graphs and semantic search
- Real-time inference and agentic AI systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a production ML inference system you've designed and scaled. What were the primary bottlenecks and how did you address them?
- How do you approach monitoring and observability for model-serving infrastructure, and what metrics matter most to you?