Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Basis Conversations 1500: 1,500 hours of multilingual, multi-party conversation

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Alexis Sursock
Word Count
779
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

Basis Conversations 1500 is a freely available dataset for commercial and research use containing 1,502 hours of multilingual, full-duplex conversations among two to four simultaneous speakers. It includes 1,907 conversations from 2,645 pseudonymous speakers across 33 countries and 22 languages, recorded as synchronized, channel-separated 48 kHz audio tracks that preserve natural variation in microphones and environments. About 100 hours have dense human annotations covering conversational nuances such as backchannels, laughter, interruptions, pauses, repairs, overlaps, and addressee intent, with machine-generated labels supplied separately where applicable. Participants were paid, matched with strangers speaking the same language, and consented to recording and licensing, while identified spoken personal details were muted and documented. Metadata, transcripts, annotations, and speaker information can be inspected before downloading the approximately 253 GB repository, though the automatically generated transcripts may be inaccurate and are not recommended as unverified training targets. Basis released the collection to support research and development in real-time human-AI interaction, including turn-taking, emotional understanding, humor, and other complex features of natural conversation.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 747 162 79 -85%
Real-time 1 649 155 80 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.