Clera
Founding AI Engineer
About this role
Build evaluation systems and feedback loops for an AI-powered B2B pricing platform. You'll own LLM infrastructure, model routing across frontier providers, and observability tools that help the AI earn trust through measurable customer revenue outcomes.
What you'll do
- Design and build evaluation harnesses and benchmarks using tracked pricing outcomes as ground truth for AI recommendations
- Own LLM infrastructure and model routing across Anthropic, Google, and other providers with explicit cost/latency/quality tradeoffs
- Maintain EU data-residency boundaries for customer model calls to ensure compliance with data regulations
- Extend MCP server infrastructure so LLM agents can programmatically drive the platform
- Build observability and feedback loops that surface where AI earns or loses trust from human reviewers
- Communicate non-deterministic engineering concepts clearly to non-technical stakeholders
What they're looking for
- 8+ years software engineering with recent production LLM experience
- LLM evaluation and observability system design
- TypeScript and Python
- FastAPI backend development
- GCP infrastructure
- Model routing and multi-provider LLM orchestration
- MCP (Model Context Protocol) development
- EU data compliance and residency requirements
Benefits
- Meaningful employee stock option plan (ESOP)
- Fully remote-friendly role based in Amsterdam with relocation support considerations
- Work on applied AI with directly measurable business impact
- High-velocity founding team environment
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a production LLM feature you shipped end-to-end—what were the biggest challenges with evaluation and observability?
- How have you approached building evals when ground truth is hard to define or noisy?