Skip to main content

Aircall.io, Inc.

Machine Learning Engineer (Evals and Voice Models)

  • Confirmed live in the last 24 hours
  • $181k–$250k
  • Mid level
  • On-site · San Francisco Office
  • 3+ yrs exp
  • Added 3 weeks ago

About this role

Aircall is seeking a Machine Learning Engineer to build and maintain robust evaluation frameworks for their rapidly growing AI-powered customer communications platform. You will focus on assessing voice models, designing automated evaluation pipelines, and ensuring consistent quality measurements across various AI products. This role requires a strong understanding of model evaluation, voice technologies, and the ability to collaborate effectively within a fast-paced, data-driven environment.

What you'll do

  • Design and implement evaluation frameworks for AI agents across multiple communication channels.
  • Train and fine-tune voice models (TTS, ASR, speech-to-speech) to improve performance.
  • Build and maintain live quality monitoring systems for deployed AI agents.
  • Develop automated regression testing and benchmarking pipelines.
  • Design annotation guidelines and calibrate LLM-as-judge systems for reliable evaluations.
  • Analyze system designs and pinpoint potential failure points.

What they're looking for

  • Machine Learning Engineering
  • Model Evaluation
  • Voice Models (TTS, ASR, Speech-to-Speech)
  • LLM Evaluation
  • Automated Evaluation Systems
  • Failure Analysis
  • Data Pipeline Construction
  • Regression Testing

Benefits

  • Competitive salary package
  • Work-life balance
  • Fast-learning environment
  • Strong team spirit
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Aircall.io, Inc.

View all jobs at Aircall.io, Inc.

Likely interview questions

  • Describe your experience building and implementing automated evaluation systems for machine learning models.
  • How do you approach failure analysis and debugging in machine learning pipelines?