Home / Companies / LangChain / Blog / Post Details
Content Deep Dive

Quickly Start Evaluating LLMs With OpenEvals

Blog post from LangChain

Post Details
Company
Date Published
Author
-
Word Count
844
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

Evaluations (evals) are essential for deploying reliable LLM-powered applications, providing systematic methods to assess LLM output quality based on specific criteria. The newly introduced packages, openevals and agentevals, offer a set of evaluators and a framework to simplify the process of building evaluations from scratch. Evals involve two components: the data being evaluated and the metrics used for evaluation, both of which significantly impact the reflection of real-world usage. The packages focus on common evaluation types, including LLM-as-a-judge evals for natural language outputs and structured data evaluations for extracting or generating structured content. Additionally, agent evaluations assess the sequence of actions taken by an agent to complete tasks. Openevals and agentevals provide tools to customize evaluations, incorporate human preferences, and ensure consistency, while LangSmith offers capabilities for tracking and sharing evaluation results. Future plans include expanding the libraries with more specific evaluators and encouraging community contributions through GitHub.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 13 4,013 569 191 -13%
Multi-agent systems 1 217 52 31 +189%
Observability 1 1,454 304 103 +17%
RAG 1 1,528 261 92 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.