Voice AI Company Gradium Now Tracked in Plushcap
July 14, 2026 by Matt Makai
Gradium is now tracked in Plushcap and has been added to two competitive spaces: Speech Understanding, Transcription, and Voice Bots and Voice Agents. Gradium publicly launched in September 2025 and was founded by researchers from Kyutai, a spin-out of research talent from Google DeepMind and Meta. Gradium announced a $70 million seed round at launch, then brought its total funding to $100 million in July 2026 with Nvidia joining as an investor. That's a lot of funding for a seed round!
The core product is a set of voice AI APIs for text-to-speech (TTS), speech-to-text (STT), speech-to-speech translation, and voice cloning. Gradium also considers Gradbot, an open-source voice agent framework, as a core product. They are targeting software developers and enterprises building production voice agents, with a secondary focus on regulated or compliance-sensitive deployments.
Voice Infrastructure with Low Latency
Gradium claims that its audio language models (ALMs) unify generation, transcription, transformation, and dialogue into one architecture to produce meaningfully lower latency than pipeline-based competitors. Their blog content hammers on this claim repeatedly in posts such as Time to First Audio, Optimizing Quality vs. Latency, and the Gradium Translate launch post. All of those posts benchmark against OpenAI Realtime, Gemini 3.5 Live, and ElevenLabs.
The Gradium on AWS post launched a managed SaaS subscription and a deployable SageMaker image for in-VPC use.
Blog Content: High Technical Density
Gradium has published 22 posts totaling ~17,800 words since launch, averaging roughly 1.8 posts per month.
A few more posts with product directions stand out:
-
Phonon (on-device TTS): The Phonon posts describe a ~100M-parameter TTS model that runs on a single CPU core, targeting offline, privacy-sensitive, and high-volume consumer deployments. The May 2026 update reports 1.00% WER on the Seed-TTS English benchmark, outperforming larger models. This is competing with embedded solutions in gaming, consumer hardware, and accessibility applications. It's currently in private beta.
-
Gradium Translate: The June 2026 launch introduced two models:
stt-translateands2s-translate. Both of these condense the three-stage translation pipeline (STT → translate → TTS) into two steps. The post benchmarks against GPT Realtime Translate and Gemini 3.5 Live Translate. The Acolad partnership for enterprise AI interpreting appears to be the first named commercial application of this capability. -
Semantic VAD: The turn detection post describes integrating turn-completion predictions directly into the STT audio model rather than relying on silence detection. Premature interruptions and sluggish responses are common failure modes for production voice AI applications, so building this capability directly into the model rather than as a post-processing step is a reasonable architectural choice for improvement in this area.
Customer Traction & Developer Community
Gradium published several partnership and customer posts: RMC BFM Drive in Renault vehicles (launched March 2026), Acolad for enterprise AI interpreting, InteractionLabs (Ongo robot), Wonderful's voice agents, and the Invincible Voice assistive technology project with Kyutai. The RMC BFM Drive case is a named deployment in a consumer product available through Renault's app store. However, none of these posts include usage metrics, contract values, or scale data. They establish that Gradium has real customers across automotive, media, robotics, and enterprise services, but the depth of those relationships is not yet visible from public sources.
Gradium has almost no Hacker News traction: only two posts have appeared, receiving six total upvotes. The YouTube channel has 84 subscribers, 12 videos, and 2,782 views, so essentially no traction there as a new channel. They have a nascent Discord but otherwise they have a long way to go to build a developer brand.
Matt's Outlook
The open question is whether Gradium’s approach of creating and owning the model layer across STT, TTS, translation, and voice AI agent orchestration provides an advantage or whether the combined package will struggle to compete with point solutions. Competitors like ElevenLabs and AssemblyAI started in individual areas like STT or TTS and used that as a way to build great products while establishing their credibility with software developers. Vapi and Retell AI are focused on the agent orchestration layer and integrated third-party models. Gradium is betting that vertical integration across the stack produces better latency and quality tradeoffs.
Whether Gradium can build the developer community needed to make Gradbot a meaningful distribution channel determines whether the company's integrated approach across models and frameworks creates enough top-of-funnel demand to sustain the high growth rates that will be required of a company that just raised $100 million in seed funding.
