Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD
Blog post from Modular
Google DeepMind has released the Gemma 4 family of models, which are state-of-the-art open multimodal models supporting text, images, and video, with enhanced performance available on both NVIDIA and AMD hardware through Modular Cloud. The Gemma 4 31B model boasts a 31-billion-parameter dense architecture with a 256K context window for complex tasks, while the Gemma 4 26B A4B is a Mixture-of-Experts model that activates only 4 billion parameters per pass to reduce compute costs. Modular Cloud offers a seamless transition from testing to production, leveraging the MAX inference framework to optimize performance and ensure consistency across different workloads. With 15% faster throughput on NVIDIA B200 compared to vLLM, Gemma 4 provides high efficiency without accuracy loss, making it one of the most capable open models available for developers and enterprises eager to deploy advanced AI solutions.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Coding Assistant | 1 | 1,480 | 382 | 153 | +18% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.