Home / Companies / LllamaIndex / Blog / Post Details
Content Deep Dive

LlamaIndex and Kaggle Launch a Document Extraction Leaderboard for AI Agents

Blog post from LllamaIndex

Post Details
Company
Date Published
Author
LlamaIndex
Word Count
1,148
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

LlamaIndex and Kaggle have launched ExtractBench, an open benchmark and leaderboard for assessing document extraction systems on 370 enterprise documents comprising 4,869 pages across eight business domains and 67 document types. Designed for workflows in which AI agents act on structured data extracted from complex files, the benchmark evaluates 14 frontier vision-language models, coding agents, and specialized APIs using deterministic, rule-based scoring across factors including extraction accuracy, OCR and perception quality, table handling, document length, domain coverage, and cost. Unlike prior benchmarks focused on short, clean digital documents, ExtractBench includes scanned, handwritten, rotated, lengthy, and table-heavy files, while using frozen schemas to enable direct comparison. Reported findings indicate that systems diverge substantially on long documents and complex tables, with LlamaIndex’s agentic extraction offering strong reported accuracy and source bounding boxes at a lower per-page cost than certain coding-agent alternatives. The initiative aims to provide transparent, reproducible evaluation through Kaggle, invite external system submissions, and expand toward harder document cases and end-to-end agent workflow testing.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 2 No monthly metrics for this publish month.
LLM 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.