Clera
Founding AI Engineer
About this role
Seed-stage wearable AI company is hiring a Founding AI Engineer to build production agentic Vision-Language Model systems running on industrial smart glasses for field technicians in data centers and energy infrastructure. You'll own the full ML stack from model training and edge optimization to evaluation frameworks and real-time voice/video interfaces.
What you'll do
- Build and ship production agentic VLM pipelines on edge devices with multi-step visual reasoning against customer workflows
- Optimize model orchestration and runtime performance for edge inference across variable connectivity conditions
- Design evaluation harnesses and data flywheels for failure capture and continuous quality improvement
- Develop real-time voice and video AI interfaces tailored to operator profiles and use cases
- Build RAG pipelines for enterprise knowledge base creation and querying from field data
- Drive multimodal model training including SFT, RL post-training, and quantization for on-premise deployment
What they're looking for
- Vision-Language Models and video-language architectures
- Production multimodal and computer vision systems
- Model orchestration and agentic AI in production
- Edge inference optimization and on-device ML
- PyTorch, vLLM, TensorRT, and model serving
- Rigorous ML evaluation methodologies and testing
- RAG, fine-tuning, RLHF, and quantization techniques
- Python and distributed training frameworks
Benefits
- Early-stage equity as founding team member
- Competitive salary with upside potential
- Work on cutting-edge wearable AI technology
- High ownership and hands-on product impact
- On-site collaborative environment in Palo Alto
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Walk us through a production multimodal or VLM system you've shipped end-to-end — how did you handle model quality, latency, and edge deployment tradeoffs?
- Describe your experience building evaluation frameworks for multimodal models. How would you measure whether a visual reasoning agent is working correctly in the field?