Introducing Ink-2: The #1-ranked STT built for voice agents
Blog post from Cartesia
Ink-2 is a cutting-edge speech-to-text model designed for real-time voice agents, excelling in accuracy, turn detection, and latency, which are critical for seamless voice interactions. It ranks #1 on Artificial Analysis’s streaming leaderboard for its low word error rate and superior built-in turn detection, enabling precise listening and response timing. The model is proficient in structured entity recognition and maintains accuracy across various accents and challenging audio conditions, outperforming competitors like Deepgram Flux and ElevenLabs Scribe v2. Ink-2 uses semantic endpointing for turn detection, allowing it to understand conversational nuances without relying solely on silence, resulting in fewer interruptions and smoother interactions. The model's latency is remarkably low, with a Time-to-Final-Transcript of 0.1 seconds, ensuring that voice agents feel responsive and attentive. Ink-2 is available via API and integrates with platforms like LiveKit, Vapi, and Pipecat, with plans for multilingual support to accommodate diverse user needs.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.