May 2025 Summaries
1 posts from Mixedbread
Filter
Month:
Year:
Post Summaries
Back to Blog
Retrieval-augmented generation systems commonly depend on OCR to convert enterprise documents into searchable text, but the reported OHR benchmark results indicate that extraction errors substantially limit both retrieval and answer quality, especially for layouts involving handwriting, charts, formulas, and complex tables. Across more than 8,500 pages and 8,498 questions, the tested OCR systems fell roughly 4.5 percentage points behind ground-truth text on NDCG@5 retrieval, while a Mixedbread multimodal vector store that retrieves directly from page images scored 0.865, about 12% above perfect-text retrieval, by using layout and visual context. In end-to-end generation tests, a conventional OCR-based pipeline achieved a 0.676 correct-answer rate, compared with 0.740 using perfect OCR and 0.912 with perfect retrieval, whereas multimodal retrieval combined with OCR text for generation reached 0.843 and recovered about 70% of the accuracy lost to OCR limitations. Direct image-based answer generation performed less reliably, suggesting that current multimodal models are more effective at visual retrieval than precise multi-document visual reading for generation. The findings support a dual-modality approach that uses visual embeddings to locate relevant pages and high-quality extracted text to supply LLM answer generation, while preserving compatibility with future vision-capable models.
May 14, 2025
3,962 words in the original blog post.