Introducing ADE-bench: measuring how AI agents perform data work
Blog post from dbt
ADE-bench, developed by Benn Stancil and dbt Labs, is a new benchmark designed to evaluate the performance of AI agents on analytics and data engineering tasks, addressing the lack of specific benchmarks in the data community. While tools like SWE-bench assess software engineering, ADE-bench uses real-world dbt projects and databases to measure how AI models tackle the complex, messy problems faced by data practitioners. Initial results reveal varied performance across models and configurations, with dbt Fusion and the Model Context Protocol (MCP) significantly improving accuracy and efficiency. The benchmark consists of dbt projects, databases, and real-world tasks and creates a sandbox environment for agents to solve presented problems. Performance is assessed by test pass rates, costs, and runtimes, with dbt Fusion showing notable gains in pass rates. ADE-bench is open-source, encouraging community contributions to enhance its relevance and effectiveness in measuring AI capabilities in data work, with ongoing efforts to improve the dbt language framework and MCP tooling.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MCP | 9 | 2,803 | 327 | 131 | -43% |
| AI Agents | 5 | 3,616 | 674 | 184 | +28% |
| LLM | 3 | 3,836 | 662 | 193 | +2% |
| AI Coding Assistant | 1 | 710 | 191 | 84 | +14% |
| Developer Experience | 1 | 413 | 204 | 87 | -9% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.