Clera
AI Engineer - Data Platform
About this role
Join a Series A AI startup as an AI Engineer on the Data Platform team to design and build scalable backend infrastructure powering autonomous incident detection and remediation. You'll work across distributed systems, performance optimization, and observability to support both cloud and on-premises deployments for enterprise customers.
What you'll do
- Design and implement scalable, resilient infrastructure for AI-driven root cause analysis and observability workflows
- Develop foundational system components ensuring efficient resource utilization and high performance at scale
- Profile and optimize backend systems to improve throughput, reduce latency, and eliminate bottlenecks
- Build and maintain internal observability stack including logs, metrics, and traces for AI agent decision-making
- Support hybrid cloud and on-premises architecture for both SaaS and enterprise deployments
- Collaborate cross-functionally to deliver infrastructure enabling real-time incident diagnosis and remediation
What they're looking for
- Backend or infrastructure engineering
- Distributed systems design and principles
- Performance profiling and optimization
- Observability platforms (Datadog, Grafana, Splunk, or similar)
- Hybrid and multi-environment infrastructure management
- High-throughput, low-latency system design
- AI/ML workload optimization
- System architecture and design
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Tell us about a time you optimized a high-throughput, low-latency system—what was the bottleneck and how did you resolve it?
- How have you designed or worked with observability systems, and what tools have you used?