Home / Companies / dbt / Blog / Post Details
Content Deep Dive

Introducing ADE-bench: measuring how AI agents perform data work

Blog post from dbt

Post Details
Company
dbt
Date Published
Author
Kathryn Chubb
Word Count
1,756
Company Posts That Month
16
Language
English
Hacker News Points
-
Post removed?
No
Summary

ADE-bench, developed by Benn Stancil and dbt Labs, is a new benchmark designed to evaluate the performance of AI agents on analytics and data engineering tasks, addressing the lack of specific benchmarks in the data community. While tools like SWE-bench assess software engineering, ADE-bench uses real-world dbt projects and databases to measure how AI models tackle the complex, messy problems faced by data practitioners. Initial results reveal varied performance across models and configurations, with dbt Fusion and the Model Context Protocol (MCP) significantly improving accuracy and efficiency. The benchmark consists of dbt projects, databases, and real-world tasks and creates a sandbox environment for agents to solve presented problems. Performance is assessed by test pass rates, costs, and runtimes, with dbt Fusion showing notable gains in pass rates. ADE-bench is open-source, encouraging community contributions to enhance its relevance and effectiveness in measuring AI capabilities in data work, with ongoing efforts to improve the dbt language framework and MCP tooling.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MCP 9 2,803 327 131 -43%
AI Agents 5 3,616 674 184 +28%
LLM 3 3,836 662 193 +2%
AI Coding Assistant 1 710 191 84 +14%
Developer Experience 1 413 204 87 -9%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.