Building a Robust Ingestion System for Any File of Any Size
Blog post from Mixedbread
Mixedbread rebuilt its file-ingestion pipeline to handle arbitrarily large and diverse user uploads, after its original approach failed on unexpectedly common gigabyte-scale text collections and hundreds-of-gigabytes video files. Its architecture separates coarse, bounded slicing from semantic parsing: a slicer divides each file into type-appropriate units such as PDF pages, video seconds, or text characters, while parsers process one slice at a time to generate meaningful search chunks at logical boundaries, scene changes, or low-energy audio regions. Small continuation states passed through a task queue allow fixed-size workers to process files sequentially without holding entire files in memory, making individual slices retryable and protecting the rest of a job from worker failures. The system addresses unreliable input-size estimates, potentially explosive rendering costs, and unsafe real-world office documents through streaming media access, page-rendering pixel limits, and lazy chunk iteration. Before ingestion, it estimates quotas using low-cost metadata such as page counts, media duration, or byte-based approximations, then reconciles totals during processing. Although the current design prioritizes semantic integrity over within-file parallelism, Mixedbread plans a future version that can process independent slices concurrently while preserving chunk quality.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.