July 2026 Summaries
2 posts from Cartesia
Filter
Month:
Year:
Post Summaries
Back to Blog
Text-to-speech (TTS) technology offers a promising way to enhance human-computer interaction by converting written text into spoken words, yet its effectiveness is challenged by the complexity of human speech nuances such as emotion, context, and linguistic correctness. Unlike a simple text-to-audio transformation, TTS systems must navigate the "one-to-many" problem, where a single text can be spoken in numerous ways depending on factors like speaker emotion and pace. Evaluation of TTS models goes beyond basic metrics like Word Error Rate (WER) to include aspects like contextual correctness, naturalness, and robustness, which are not fully captured by automated quality scores. Human evaluations remain the gold standard but are subject to variability based on listener background and presentation details. Long-form and streaming scenarios further complicate evaluations by exposing issues like speaker consistency and streaming artifacts. Additionally, domain-specific texts and multilingual challenges require TTS systems to adapt to specialized vocabularies and pronunciation rules, highlighting the need for evaluations that are tailored to the specific use cases and linguistic contexts of the product.
Jul 28, 2026
3,501 words in the original blog post.
Ink-2 is a cutting-edge speech-to-text model designed for real-time voice agents, excelling in accuracy, turn detection, and latency, which are critical for seamless voice interactions. It ranks #1 on Artificial Analysis’s streaming leaderboard for its low word error rate and superior built-in turn detection, enabling precise listening and response timing. The model is proficient in structured entity recognition and maintains accuracy across various accents and challenging audio conditions, outperforming competitors like Deepgram Flux and ElevenLabs Scribe v2. Ink-2 uses semantic endpointing for turn detection, allowing it to understand conversational nuances without relying solely on silence, resulting in fewer interruptions and smoother interactions. The model's latency is remarkably low, with a Time-to-Final-Transcript of 0.1 seconds, ensuring that voice agents feel responsive and attentive. Ink-2 is available via API and integrates with platforms like LiveKit, Vapi, and Pipecat, with plans for multilingual support to accommodate diverse user needs.
Jul 09, 2026
804 words in the original blog post.