Mercury 2, the first reasoning diffusion LLM, is now on Baseten
Blog post from Baseten
Inception's Mercury 2, now available on Baseten, is the first inference platform to offer production-grade diffusion LLMs (dLLMs), providing a significant speed and cost advantage over traditional autoregressive models. It achieves over 1,000 tokens per second on standard NVIDIA GPUs, without the need for specialized hardware, making it both fast and economically viable for developers. Mercury 2's innovative architecture allows for parallel token drafting and refinement, enhancing speed and opening new application possibilities in areas like real-time voice agents and efficient coding tools. Notably, Augment Code has successfully implemented Mercury 2, reducing costs by 90% and latency by 82% for critical tasks. Baseten supports Mercury 2 with enterprise-grade infrastructure, ensuring reliable, scalable, and compliant deployment, which allows Inception to focus on optimizing AI workflows while maintaining efficiency.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 6 | 6,237 | 1,165 | 246 | -31% |
| Observability | 2 | 4,230 | 776 | 198 | +24% |
| Real-time | 2 | 5,758 | 1,361 | 266 | +0% |
| AI Coding Assistant | 1 | 2,161 | 541 | 167 | +20% |
| MCP | 1 | 7,668 | 844 | 209 | +8% |
| Multi-agent systems | 1 | 538 | 169 | 80 | -1% |
| Voice AI | 1 | 3,155 | 274 | 58 | -9% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.