Home / Companies / LllamaIndex / Blog / Post Details
Content Deep Dive

Introducing ExtractBench: The Most Comprehensive Benchmark for Data Extraction from Enterprise Documents

Blog post from LllamaIndex

Post Details
Company
Date Published
Author
LlamaIndex
Word Count
1,963
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

ExtractBench is an open, reproducible benchmark for evaluating enterprise document-extraction systems under production-oriented conditions, including long records, scanned and handwritten pages, complex tables, evidence grounding, and per-page cost. Its corpus contains 370 real and synthetic enterprise documents totaling 4,869 pages across eight business domains and 67 document types, with schemas tailored to each type and ground truth created through cross-system review, data-first synthetic generation, and manual form annotation. The benchmark evaluates 14 frontier vision-language models, coding agents, and specialized APIs, finding that many systems perform well on short documents but lose recall substantially on long documents, scans, handwriting, or complex structural tasks. LlamaExtract Agentic Plus ranked highest overall with a 95.6% value F1 score at 8.1 cents per page, retained 94.4% accuracy on the longest documents, and led systems that provide grounding evidence, while the report notes that precise value-level grounding remains unresolved across the field. The authors emphasize that price does not reliably predict accuracy and make the dataset, evaluation harness, schemas, and methodology publicly available for independent testing.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.