Skip to main content

Clera

ML Infrastructure Engineer

San MateofulltimemidAdded today

About this role

Early-stage enterprise AI startup seeks an ML Infrastructure Engineer to design and operate production inference and model-serving systems that power AI agents in regulated industries. You'll own the full lifecycle of inference infrastructure, optimize for reliability and scale, and collaborate across ML and platform teams to solve infrastructure challenges in real-world deployments.

What you'll do

  • Design, build, and operate end-to-end inference and model-serving infrastructure for production AI agents
  • Scale systems to handle increasing concurrency while maintaining reliability and performance
  • Collaborate with ML and infrastructure teams to integrate and optimize model-serving platforms
  • Identify and resolve infrastructure bottlenecks through cross-functional engineering solutions
  • Optimize production systems for latency, throughput, and resource efficiency
  • Implement observability and monitoring to ensure system reliability in production

What they're looking for

  • ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
  • Kubernetes and Docker containerization for ML workloads
  • Distributed systems design with concurrent request handling
  • Observability tools (Prometheus, Grafana, ELK, distributed tracing)
  • Cloud platforms (AWS, GCP, or Azure)
  • Systems or backend languages (Python, Go, Rust, C++, or Java)
  • Production ML system optimization for latency and throughput
  • Knowledge graphs or semantic search (nice-to-have)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Describe a production inference system you scaled under heavy concurrent load—what were the bottlenecks and how did you resolve them?
  • How have you optimized latency and throughput in a model-serving platform, and which framework did you use?