Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

FineBooks: are open OCR models good enough to unlock historical knowledge?

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Sebastian Majstorovic and Daniel van Strien
Word Count
2,151
Company Posts That Month
52
Language
-
Hacker News Points
-
Post removed?
No
Summary

FineBooks, a collaboration between Hugging Face and EleutherAI, has launched an open benchmark assessing whether modern open-weight OCR models can accurately and affordably reprocess historical books, whose existing digitized text often reflects outdated OCR systems. Using 2,165 expert-corrected pages from six Biodiversity Heritage Library volumes in English, French, German, and Latin, the project evaluated 14 permissively licensed models through reproducible Hugging Face Jobs and released the underlying ground-truth dataset and evaluation tools. Results measured character error rate, recall, over-extraction, and repetitive-output failures, with dots.mocr achieving 97.6% reading accuracy and other leading models approaching similar performance at costs ranging from cents to a few dollars per thousand pages. The findings suggest that current models are generally suitable for improving large open training corpora such as Common Pile, but are less straightforward for library workflows dependent on word-level ALTO XML coordinates and remain insufficient for scholarly transcription because they often modernize historical characters such as the long s and ligatures. The benchmark is limited to antiqua-family printed books and does not assess Fraktur, handwriting, non-Latin scripts, or more complex materials such as newspapers, while FineBooks plans to continually add models and re-OCR approximately 200,000 public-domain BHL items for release as open data.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 2,482 499 155 -67%
AI Model Fine-tuning 1 278 80 43 -70%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.