Home / Companies / LangChain / Blog / Post Details
Content Deep Dive

How We Benchmark Deep Agents

Blog post from LangChain

Post Details
Company
Date Published
Author
Nick Hollon, Harrison Chase
Word Count
712
Company Posts That Month
23
Language
English
Hacker News Points
-
Post removed?
No
Summary

Agent design challenges are compounded by the difficulty of evaluation, prompting the development of Deep Agents, an open-source, model-agnostic agent harness. The team behind Deep Agents revamped their evaluation framework, moving from smaller unit tests to comprehensive end-to-end evaluations using Harbor, an open-source framework known for powering Terminal Bench. These evaluations require an environment, instruction, and evaluation script, distinguishing them from simpler LLM evaluations. Three benchmarks—Harbor-Index, 𝜏³-bench, and ContextBench—cover various agent tasks across domains like software engineering and data analysis. The team employs practices such as running tasks multiple times to account for nondeterminism and maintaining a "lite" benchmark for rapid iteration. These benchmarks guide decision-making and iteration, exemplified by preparations for a 0.7 release of Deep Agents, where the team considers removing unnecessary components like the todo-list middleware to optimize performance.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.