Introducing LangSmith Tuned Evaluators, starting with Perceived Error
Blog post from LangChain
LangChain has introduced Tuned Evaluators for LangSmith, managed and versioned evaluators that automatically assess production agent traces and threads, beginning with an evaluator for Perceived Error. Designed to identify evidence that an agent made mistakes, misunderstood users, repeated unresolved behavior, or otherwise led an interaction in the wrong direction, Perceived Error provides a proxy for user satisfaction when explicit ratings are unavailable. LangChain manages the evaluator prompts, specialized judge models, benchmarking, infrastructure, and maintenance, allowing teams to attach an evaluator to a tracing project and receive feedback and explanations on eligible conversations. The company states that its post-trained model outperformed frontier models in its benchmark while reducing evaluation costs by 82%, with some early users reporting larger savings. Teams can use flagged traces to investigate recurring failures, create evaluation datasets, route ambiguous cases for human review, and validate agent improvements. Perceived Error evaluations become eligible after at least two human-AI message pairs and an idle period, complete within 12 hours, and are currently available to LangSmith Plus and Cloud Enterprise customers in the United States.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 2 | 3,175 | 737 | 186 | -24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.