LangChain
Research Engineer, LangSmith Engine
About this role
Research Engineer role focused on improving LangSmith Engine, an autonomous agent that analyzes production traces and recommends fixes. You'll build benchmarks, run experiments to enhance agent performance across models and strategies, and transition successful research into production systems that balance quality, cost, and latency.
What you'll do
- Build and maintain benchmarks and evaluations measuring agent quality and efficiency on real-world tasks
- Design and run experiments to improve agent performance across models, prompting, context, tools, and orchestration strategies
- Explore and implement post-training and fine-tuning techniques to enhance agent capabilities, quality, or cost
- Convert successful experiments into production improvements while measuring impact and preventing regressions
- Define ML roadmap and technical direction for Engine agents, providing technical leadership and mentorship
- Balance research improvements against production constraints including cost, latency, reliability, and scalability
What they're looking for
- Machine learning research and experimentation design
- Large language models and AI agent development
- Benchmark and evaluation framework design
- Post-training techniques (fine-tuning, RLHF, preference optimization)
- Software engineering and production systems thinking
- Prompting optimization and model selection
- Data analysis and statistical rigor
- Cross-team collaboration and technical communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
LangChain
LangChain builds platforms and frameworks for developing, deploying, and observing production AI agents at enterprise scale, including LangSmith for AI observability and evaluation. The company is hiring Deployed Engineers to work directly with enterprise customers on agent implementation and operations, as well as Fullstack Engineers to build features across its platform stack.
- Website
- langchain.com
Likely interview questions
- Describe your experience working with LLMs in production—what was the most challenging performance issue you debugged and how did you approach it?
- Walk us through how you would design a benchmark to evaluate an autonomous agent system. What metrics matter most and why?