Home / Companies / Cartesia / Blog / October 2024

October 2024 Summaries

2 posts from Cartesia

Filter
Month: Year:
Post Summaries Back to Blog
Voice Changer is a newly launched model that transforms audio voices while preserving the original delivery and emotion, offering a diverse library of studio-quality voices or the ability to clone one's own voice with precise control over vocal nuances such as vocalization, emotion, and prosody. It demonstrates seamless voice transitions and allows for creative flexibility in various applications, including content creation, gaming, entertainment, and business, by enabling users to modify and generate high-quality audio that aligns with their needs. Built on advanced state space model architectures like S4 and Mamba, Voice Changer supports high-resolution audio processing, making it efficient for generating realistic voice transformations. The platform provides an accessible playground for users to experiment and invites developers to explore innovative applications, promising enterprise support and ongoing AI advancements through hiring and community engagement.
Oct 30, 2024 537 words in the original blog post.
Sonic Multilingual has launched its Alpha Release, expanding its language support to a total of 15 by adding Hindi, Italian, Korean, Dutch, Polish, Russian, Swedish, and Turkish. This enhancement enables developers and users to create lifelike speech applications across a diverse range of languages and regions, featuring the same industry-leading latency and conversational quality as Sonic English. The platform is based on an efficient state space model (SSM) architecture, allowing real-time, high-resolution audio generation with near-linear scaling costs. Users can easily select the "sonic-multilingual" model in their applications, choose from 14 available languages, type a transcript, pick a voice, and experience low-latency voice generation. Sonic encourages partnerships to build real-time conversational AI and invites interested individuals to join their team as they expand their multilingual and multimodal intelligence capabilities.
Oct 08, 2024 533 words in the original blog post.