Home / Companies / Weaviate / Blog / Post Details
Content Deep Dive

Ingesting PDFs into Weaviate

Blog post from Weaviate

Post Details
Company
Date Published
Author
Erika Cardenas, Mohd Shukri Hasan
Word Count
1,776
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

The latest advancements in multimodal deep learning have made it possible to extract high quality data from PDF documents and add it to a Weaviate workflow. Optical Character Recognition (OCR) technology is used to convert different types of visual documents into machine-readable formats, with new models like LayoutLMv3 and Donut leveraging both text and visual information using multimodal transformers. Unstructured, an open-source company working at the cutting edge of PDF processing, allows businesses to ingest diverse data sources and convert them into data that can be passed to a Language Learning Model (LLM). This enables users to chat with their PDFs by converting private documents from their company into text format.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 3 1,416 172 75 +112%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.