Basis Conversations 1500: 1,500 hours of multilingual, multi-party conversation
Blog post from Hugging Face
Basis Conversations 1500 is a freely available dataset for commercial and research use containing 1,502 hours of multilingual, full-duplex conversations among two to four simultaneous speakers. It includes 1,907 conversations from 2,645 pseudonymous speakers across 33 countries and 22 languages, recorded as synchronized, channel-separated 48 kHz audio tracks that preserve natural variation in microphones and environments. About 100 hours have dense human annotations covering conversational nuances such as backchannels, laughter, interruptions, pauses, repairs, overlaps, and addressee intent, with machine-generated labels supplied separately where applicable. Participants were paid, matched with strangers speaking the same language, and consented to recording and licensing, while identified spoken personal details were muted and documented. Metadata, transcripts, annotations, and speaker information can be inspected before downloading the approximately 253 GB repository, though the automatically generated transcripts may be inaccurate and are not recommended as unverified training targets. Basis released the collection to support research and development in real-time human-AI interaction, including turn-taking, emotional understanding, humor, and other complex features of natural conversation.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.