Home / Companies / Encord / Blog / Post Details
Content Deep Dive

How We Built the World's Largest Multimodal Dataset

Blog post from Encord

Post Details
Company
Date Published
Author
Frederik Hvilshøj
Word Count
2,211
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

Over the past few months, the Encord machine learning team has developed what they claim to be the world's largest open-source multimodal dataset, designed to support the development of models that integrate text, images, video, audio, and 3D point clouds. This dataset aims to facilitate advancements in multimodal AI by providing a clean and extensive resource for open-source development. The process involved sourcing data from multiple modalities, using retrieval models to align the data, and enhancing data quality through human annotation. They also created a retrieval model capable of embedding all modalities into a common space, evaluated through public benchmarks and a newly built dataset for audio-point cloud embeddings. A baseline retrieval model was trained, demonstrating that high-quality data can outperform larger parameter models in cross-modal retrieval tasks. The Encord team hopes that sharing their methodology will aid others in constructing similar datasets and furthering multimodal AI innovation.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 15 1,589 336 137 +6%
LLM 4 4,863 783 205 +34%
AI Guardrails 1 285 103 50 -30%
AI Model Fine-tuning 1 762 158 56 +176%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.