Home / Companies / Cartesia / Blog / March 2025

March 2025 Summaries

6 posts from Cartesia

Filter
Month: Year:
Post Summaries Back to Blog
The release of version 2.0.0 of Cartesia's Python SDK enhances the developer experience for utilizing Cartesia's AI voice capabilities in Python. The SDK is designed around a primary Cartesia client, serving as the main access point for various API endpoints, and includes features such as a basic client, an async client for non-blocking requests, and streaming support. It also offers integration with WebSockets for real-time, low-latency applications and incorporates robust exception handling, automatic retries with exponential backoff, and customizable timeouts, alongside an option to override the httpx client for specific use cases like proxy support. This release encourages developers to explore its features and provide feedback, emphasizing its potential for building real-time, multimodal intelligence for devices, and is accompanied by new tools like Ink for speech-to-text models and organizational dashboards.
Mar 28, 2025 410 words in the original blog post.
Cartesia, a prominent developer of generative audio models, has been recognized in the seventh annual Enterprise Tech 30, a prestigious list identifying the most promising private enterprise tech companies poised to revolutionize business operations. This accolade highlights Cartesia's potential for significant impact and positions it among startups expected to become major IPOs or achieve multi-billion dollar exits. The selection process for the Enterprise Tech 30 involves a stringent evaluation by over 100 venture capitalists, considering nearly 600 venture-backed companies. Companies are categorized based on their total capital raised, with stages ranging from early to Giga. The list, founded by Wing Venture Capital's Peter Wagner, is a testament to a company's influence and future potential in enterprise technology, and being named is a significant achievement that underscores a company's role in driving business transformation.
Mar 25, 2025 288 words in the original blog post.
Narrations is a cutting-edge platform designed to transform written content into high-quality audio productions, offering creators exceptional control and efficiency. It caters to various forms of audio content like audiobooks, podcasts, and narrative works by providing an extensive Voice Library and Voice Design features that allow users to select from hundreds of voices and customize every aspect of the audio delivery. With capabilities that include creating character voices in over 15 languages, adjusting emotional delivery and pacing, and editing specific audio fragments, Narrations goes beyond typical text-to-speech technology by incorporating advanced features like instant voice cloning from just three seconds of audio. It also supports importing documents from popular formats and platforms, allowing for seamless integration and expressive audio generation. Powered by Sonic 2.0, Narrations represents a reimagining of audio content creation, promising users a futuristic and immersive experience with AI-driven voice capabilities.
Mar 18, 2025 246 words in the original blog post.
Cartesia is advancing the field of voice AI with the launch of its latest model, Sonic 2.0, following a successful $64 million Series A funding round led by Kleiner Perkins. The company has developed a voice generation model that is faster and more controllable than its predecessor, offering twice the size yet improved speed and is preferred by users over other providers. Sonic 2.0 enhances voice cloning capabilities, capturing complex accents and soundscapes, and introduces new features like 'Voice Changer' and 'Infill' for audio customization and editing. Cartesia's platform, designed for developers, boasts enterprise-grade infrastructure with high reliability and compliance standards, supporting real-time deployments. The company is committed to ongoing research to advance future audio models, focusing on areas like streaming architectures and on-device inference.
Mar 11, 2025 337 words in the original blog post.
The text discusses the shift towards on-device AI, emphasizing the need for efficient models that can operate across various hardware environments to support applications like personal assistants and real-time translators. The research introduces "Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing," which explores architecture distillation—a method to transform pre-trained models into more efficient architectures like Mamba-2, enhancing inference performance while maintaining model quality. The paper highlights the benefits of this approach, including efficiency gains, deployment flexibility, and the advancement of small model capabilities. A new distillation framework, MOHAWK, is introduced to convert Transformer models into efficient Mamba-2 variants with significantly less data and compute than traditional methods. The research demonstrates how optimized Mamba-2 models, integrated with Apple's Metal framework, deliver high throughput and reduced memory usage, with Llamba models achieving up to 12X higher token processing throughput compared to their Transformer-based counterparts. This innovation is poised to facilitate the decentralization of AI, enabling responsive and accessible AI applications across devices.
Mar 05, 2025 1,063 words in the original blog post.
The text discusses the advancements in voice AI technology, highlighting the shift from traditional voice assistants to more sophisticated AI agents capable of engaging in natural, context-aware conversations. These agents enhance customer experiences by using speech recognition, natural language processing, and text-to-speech technology to perform tasks like answering questions, scheduling appointments, and managing calls. The Cartesia API is presented as a tool for building these voice AI systems, offering customization options such as speed and emotion adjustments to create realistic conversational experiences. The text also mentions the potential applications of voice AI across various industries, noting their ability to support multiple languages and integrate seamlessly with existing systems. Furthermore, it emphasizes that modern AI-generated speech has improved significantly, reducing the robotic sound and making it indistinguishable from human speech.
Mar 03, 2025 1,896 words in the original blog post.