EMA
AI Application Engineer (Contractor role)
About this role
Ema seeks an AI Application Engineer to measure and improve the performance of healthcare AI agents in production. You'll instrument agent behavior, identify quality issues, and design experiments to enhance accuracy and clinical soundness in real-world healthcare workflows.
What you'll do
- Instrument and monitor healthcare AI agent performance metrics in production environments
- Identify failure modes and quality breakdowns in agent outputs
- Design and run experiments to validate improvements and measure impact
- Collaborate with clinical experts to ensure improvements drive better patient outcomes
- Analyze data to inform agent optimization and refinement strategies
- Document findings and communicate results to cross-functional teams
What they're looking for
- Python programming
- Agent development and debugging
- Data analysis and experimentation design
- Metrics definition and instrumentation
- Healthcare domain knowledge (valued but not required)
- Comfort with ambiguity and incomplete information
- Technical communication with clinical stakeholders
Benefits
- Work on frontier agentic AI in healthcare
- Collaborate with founders from Google, Coinbase, Flipkart, and Okta
- Access to Bay Area office for weekly syncs
- Exposure to production AI systems at scale
- Equity eligibility for certain roles
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
EMA
EMA is an agentic AI platform that enables enterprise automation and AI-driven business processes through scalable APIs and user-facing applications. The company is hiring Full-Stack, Frontend, and Backend Software Engineers, Machine Learning Engineers, and DevOps Engineers to build and deploy production-scale AI systems.
View all jobs at EMALikely interview questions
- Describe a time you identified a bug or performance issue in a system and how you validated your fix—what metrics did you use?
- How would you design an experiment to measure whether an improvement to an AI agent actually translates to better outcomes?