Skip to main content

Clera

ML Infrastructure Engineer

San MateofulltimemidAdded today

About this role

Own end-to-end ML inference and model-serving infrastructure at an early-stage enterprise AI company, building systems that keep AI agents running reliably and securely in production. You'll design, scale, and optimize inference systems handling high concurrency while collaborating closely with ML and infrastructure teams.

What you'll do

  • Design and own inference and model-serving infrastructure from architecture through production deployment
  • Build and scale systems enabling AI agents to run reliably under high concurrency in production
  • Collaborate with ML and infrastructure teams on integration and performance optimization
  • Identify and resolve infrastructure bottlenecks and scaling challenges
  • Monitor, observe, and debug production systems for reliability and performance

What they're looking for

  • ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
  • Kubernetes and Docker containerization for ML workloads
  • Production system optimization for latency, throughput, and reliability
  • Monitoring and observability tools (Prometheus, Grafana, ELK, distributed tracing)
  • Cloud platforms (AWS, GCP, or Azure)
  • Systems or backend languages (Python, Go, Rust, C++, Java)
  • Distributed systems design and resource allocation under load
  • Knowledge graphs, semantic search, or graph databases (preferred)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Walk us through a production ML inference system you built from scratch—what were the key architectural decisions and bottlenecks you encountered?
  • How have you optimized model-serving systems for latency and throughput at scale, and what metrics did you track?