FineBooks: are open OCR models good enough to unlock historical knowledge?
Blog post from Hugging Face
FineBooks, a collaboration between Hugging Face and EleutherAI, has launched an open benchmark assessing whether modern open-weight OCR models can accurately and affordably reprocess historical books, whose existing digitized text often reflects outdated OCR systems. Using 2,165 expert-corrected pages from six Biodiversity Heritage Library volumes in English, French, German, and Latin, the project evaluated 14 permissively licensed models through reproducible Hugging Face Jobs and released the underlying ground-truth dataset and evaluation tools. Results measured character error rate, recall, over-extraction, and repetitive-output failures, with dots.mocr achieving 97.6% reading accuracy and other leading models approaching similar performance at costs ranging from cents to a few dollars per thousand pages. The findings suggest that current models are generally suitable for improving large open training corpora such as Common Pile, but are less straightforward for library workflows dependent on word-level ALTO XML coordinates and remain insufficient for scholarly transcription because they often modernize historical characters such as the long s and ligatures. The benchmark is limited to antiqua-family printed books and does not assess Fraktur, handwriting, non-Latin scripts, or more complex materials such as newspapers, while FineBooks plans to continually add models and re-OCR approximately 200,000 public-domain BHL items for release as open data.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 2,482 | 499 | 155 | -67% |
| AI Model Fine-tuning | 1 | 278 | 80 | 43 | -70% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.