Skip to main content

IFM

Eval360 - Error Analysis Engineer

Sunnyvale, CA$150k–$450kfull-timemidAdded 1 month ago

About this role

Join the Institute of Foundation Models as an Error Analysis Engineer to develop systems that evaluate and improve foundation models. You'll work with a multidisciplinary team to identify failure modes, measure model quality, and enhance the safety and reliability of AI systems.

What you'll do

  • Analyze errors and failure modes in foundation models
  • Develop evaluation systems to measure model quality and performance
  • Collaborate with researchers and ML engineers on model improvement strategies
  • Support risk assessment and safety evaluation of frontier models
  • Contribute to model governance and deployment readiness frameworks
  • Work with product teams to translate research into impactful systems

What they're looking for

  • Error analysis and debugging methodologies
  • Machine learning evaluation techniques
  • Data analysis and statistical reasoning
  • Python or similar programming languages
  • Foundation models knowledge
  • Research and problem-solving abilities
  • Cross-functional collaboration
  • AI safety and model assessment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

IFM

Institute of Foundation Models conducts research on large-scale foundation models, diffusion-based language models, and world models, building the distributed training infrastructure and MLOps systems required for cutting-edge AI development. The company is hiring research scientists, ML infrastructure engineers, and systems developers to optimize pre-training frameworks, scale distributed training across multi-GPU clusters, and advance inference and experiment capabilities.

View all jobs at IFM

Likely interview questions

  • Can you describe your experience with error analysis or failure mode identification in machine learning models?
  • How have you approached systematizing and categorizing different types of model errors or behavioral issues you've observed?