Home / Companies / LangChain / Blog / Post Details
Content Deep Dive

Introducing LangSmith Tuned Evaluators, starting with Perceived Error

Blog post from LangChain

Post Details
Company
Date Published
Author
Jake Broekhuizen, Shamik Karkhanis, Vivek Trivedy
Word Count
827
Company Posts That Month
19
Language
English
Hacker News Points
-
Post removed?
No
Summary

LangChain has introduced Tuned Evaluators for LangSmith, managed and versioned evaluators that automatically assess production agent traces and threads, beginning with an evaluator for Perceived Error. Designed to identify evidence that an agent made mistakes, misunderstood users, repeated unresolved behavior, or otherwise led an interaction in the wrong direction, Perceived Error provides a proxy for user satisfaction when explicit ratings are unavailable. LangChain manages the evaluator prompts, specialized judge models, benchmarking, infrastructure, and maintenance, allowing teams to attach an evaluator to a tracing project and receive feedback and explanations on eligible conversations. The company states that its post-trained model outperformed frontier models in its benchmark while reducing evaluation costs by 82%, with some early users reporting larger savings. Teams can use flagged traces to investigate recurring failures, create evaluation datasets, route ambiguous cases for human review, and validate agent improvements. Perceived Error evaluations become eligible after at least two human-AI message pairs and an idle period, complete within 12 hours, and are currently available to LangSmith Plus and Cloud Enterprise customers in the United States.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Observability 2 3,175 737 186 -24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.