Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Mercury 2, the first reasoning diffusion LLM, is now on Baseten

Blog post from Baseten

Post Details
Company
Date Published
Author
Fred Liu
Word Count
1,255
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

Inception's Mercury 2, now available on Baseten, is the first inference platform to offer production-grade diffusion LLMs (dLLMs), providing a significant speed and cost advantage over traditional autoregressive models. It achieves over 1,000 tokens per second on standard NVIDIA GPUs, without the need for specialized hardware, making it both fast and economically viable for developers. Mercury 2's innovative architecture allows for parallel token drafting and refinement, enhancing speed and opening new application possibilities in areas like real-time voice agents and efficient coding tools. Notably, Augment Code has successfully implemented Mercury 2, reducing costs by 90% and latency by 82% for critical tasks. Baseten supports Mercury 2 with enterprise-grade infrastructure, ensuring reliable, scalable, and compliant deployment, which allows Inception to focus on optimizing AI workflows while maintaining efficiency.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 6 6,292 1,205 252 -36%
Observability 2 4,261 791 201 +16%
Real-time 2 6,055 1,444 270 -11%
AI Coding Assistant 1 2,234 577 171 +12%
MCP 1 7,755 862 214 0%
Multi-agent systems 1 556 175 81 -7%
Subagents 1 350 107 61 +39%
Voice AI 1 3,175 278 59 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.