Home / Companies / Modular / Blog / Post Details
Content Deep Dive

Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD

Blog post from Modular

Post Details
Company
Date Published
Author
Modular Team
Word Count
710
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

Google DeepMind has released the Gemma 4 family of models, which are state-of-the-art open multimodal models supporting text, images, and video, with enhanced performance available on both NVIDIA and AMD hardware through Modular Cloud. The Gemma 4 31B model boasts a 31-billion-parameter dense architecture with a 256K context window for complex tasks, while the Gemma 4 26B A4B is a Mixture-of-Experts model that activates only 4 billion parameters per pass to reduce compute costs. Modular Cloud offers a seamless transition from testing to production, leveraging the MAX inference framework to optimize performance and ensure consistency across different workloads. With 15% faster throughput on NVIDIA B200 compared to vLLM, Gemma 4 provides high efficiency without accuracy loss, making it one of the most capable open models available for developers and enterprises eager to deploy advanced AI solutions.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Coding Assistant 1 1,480 382 153 +18%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.