Skip to main content

LangChain

Agent Reliability Engineer, GTM

San Francisco, CA$150k–$190kfulltimemidAdded 5 days ago

About this role

LangChain is seeking an Agent Reliability Engineer to own the health, performance, cost, and business impact of their production GTM agent. You'll monitor production systems, triage issues, run evaluations, track metrics, and build the monitoring and documentation that demonstrates the agent's value to customers and leadership.

What you'll do

  • Monitor production agent health, catching errors, slow runs, expensive runs, and silent failures
  • Triage and resolve issues from Slack, tickets, and rep reports, escalating complex problems appropriately
  • Run weekly eval suites, investigate failures, and convert production bugs into permanent regression tests
  • Analyze cost and latency by model, graph, use case, and role to recommend optimization changes
  • Track usage and adoption metrics per rep and feature, producing weekly health reports
  • Build business impact metrics showing agent value (reply rates, meetings booked, hours reclaimed, ROI)

What they're looking for

  • Production Python and SQL
  • LLM application operations (tracing, evals, prompting, caching)
  • SRE mindset (percentiles, SLOs, signal vs. noise)
  • Metrics analysis and validation
  • Technical writing and documentation
  • High agency and initiative
  • BigQuery or dbt (nice to have)
  • LangGraph or LangSmith experience (nice to have)

Benefits

  • Competitive base salary ($150,000–$190,000)
  • Variable compensation for relevant roles
  • Meaningful equity
  • Benefits and perks
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

LangChain

LangChain builds platforms and frameworks for developing, deploying, and observing production AI agents at enterprise scale, including LangSmith for AI observability and evaluation. The company is hiring Deployed Engineers to work directly with enterprise customers on agent implementation and operations, as well as Fullstack Engineers to build features across its platform stack.

View all jobs at LangChain

Likely interview questions

  • Describe a time you caught a production issue before users reported it—what signals did you monitor?
  • How would you design an eval suite to catch regressions in an LLM-powered agent?