LlamaIndex and Kaggle Launch a Document Extraction Leaderboard for AI Agents
Blog post from LllamaIndex
LlamaIndex and Kaggle have launched ExtractBench, an open benchmark and leaderboard for assessing document extraction systems on 370 enterprise documents comprising 4,869 pages across eight business domains and 67 document types. Designed for workflows in which AI agents act on structured data extracted from complex files, the benchmark evaluates 14 frontier vision-language models, coding agents, and specialized APIs using deterministic, rule-based scoring across factors including extraction accuracy, OCR and perception quality, table handling, document length, domain coverage, and cost. Unlike prior benchmarks focused on short, clean digital documents, ExtractBench includes scanned, handwritten, rotated, lengthy, and table-heavy files, while using frozen schemas to enable direct comparison. Reported findings indicate that systems diverge substantially on long documents and complex tables, with LlamaIndex’s agentic extraction offering strong reported accuracy and source bounding boxes at a lower per-page cost than certain coding-agent alternatives. The initiative aims to provide transparent, reproducible evaluation through Kaggle, invite external system submissions, and expand toward harder document cases and end-to-end agent workflow testing.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.