Skip to main content

Clera

ML Infrastructure Engineer

San MateofulltimemidAdded 2 days ago

About this role

Early-stage enterprise AI company seeks an ML Infrastructure Engineer to design, build, and operate inference and model-serving infrastructure for AI agents in regulated industries. You'll own the complete infrastructure lifecycle from development to production, ensuring reliability and performance at scale.

What you'll do

  • Design and build inference and model-serving infrastructure for production AI agent deployments
  • Scale systems to handle increasing concurrency and production load reliably
  • Identify and resolve infrastructure bottlenecks in collaboration with ML and platform teams
  • Optimize systems for latency, throughput, and reliability
  • Operate and maintain ML inference platforms through development and deployment phases
  • Monitor and debug production systems using observability tools

What they're looking for

  • ML inference systems and model-serving platforms (TensorFlow Serving, TorchServe, Triton, KServe)
  • Distributed systems design and optimization
  • Kubernetes and Docker containerization
  • Cloud infrastructure (AWS, GCP, or Azure)
  • Monitoring and observability tools (Prometheus, Grafana, ELK, distributed tracing)
  • Systems programming languages (Python, Go, Rust, C++, or Java)
  • Production ML system optimization for latency and throughput
  • Knowledge graphs or graph databases (bonus)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Clera

Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.

View all jobs at Clera

Likely interview questions

  • Describe your experience designing and scaling inference serving infrastructure in production. What challenges did you face and how did you resolve them?
  • Tell us about a time you optimized an ML system for latency or throughput under high concurrency. What metrics did you track and what was the outcome?