Skip to main content

Benchling

Software Engineer, Model Evaluation and Improvement

San Francisco, CA$136.4k–$166.8kfulltimemidAdded today

About this role

Join Benchling's Model Evaluation and Improvement team to build datasets, evaluations, and systems that enhance frontier AI models for scientific applications. You'll work at the intersection of software engineering, biology, and AI to close the gap between LLM capabilities and real-world scientific problem-solving.

What you'll do

  • Create high-quality datasets by transforming complex scientific data into tasks and evaluation environments for LLMs
  • Analyze model failure modes across frontier models to identify improvement opportunities
  • Design and implement scalable data pipelines for curating, transforming, and validating scientific data
  • Partner with frontier AI labs to develop and evaluate approaches for improving model performance on scientific tasks
  • Translate scientific expert judgment into problems and evaluation criteria that distinguish strong model behavior

What they're looking for

  • LLM development and evaluation experience
  • Biology or biotech domain knowledge
  • Data pipeline and infrastructure design
  • Scientific data curation and validation
  • Model analysis and failure mode identification
  • Cross-functional collaboration with scientists and engineers
  • Rapid experimentation and prototyping
  • Comfort with ambiguous, evolving technical problems
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Benchling

Benchling builds an AI-powered platform for biotech R&D that integrates scientific workflows and data processes to accelerate research breakthroughs. The company is hiring software engineers across full-stack, customer engineering, agentic AI, and security roles to enhance developer productivity, build production AI systems, and protect sensitive research data.

View all jobs at Benchling

Likely interview questions

  • Can you walk us through a specific example where you identified and debugged an LLM failure mode in a scientific context?
  • How have you approached designing evaluation benchmarks that capture nuanced scientific judgment?