Nebius Token Factory Becomes First AI Cloud to Adopt NVIDIA Groq 3 LPX
Blog post from Nebius
Nebius has announced that its Token Factory inference platform will be the first AI cloud to adopt NVIDIA Groq 3 LPX alongside NVIDIA Vera Rubin NVL72, targeting faster token generation for latency-sensitive agentic workloads. The company argues that as agents make many sequential model calls, generation speed increasingly determines overall task completion time, particularly when processing long contexts and maintaining multi-turn state. Each Groq 3 LPX rack combines 256 LPU accelerators, 128 GB of on-chip SRAM, and 640 TB/s of scale-up bandwidth in a liquid-cooled NVIDIA MGX architecture. Artificial Analysis reportedly measured Groq 3 LPX running Gemma 4 31B at 3,400 output tokens per second for one user, while NVIDIA projects Vera Rubin NVL72 could deliver up to 35 times greater inference throughput per megawatt than GB200 NVL72 for long-context, low-latency 2T-parameter workloads. Nebius plans to provide the new hardware through existing Token Factory capabilities, including serverless and dedicated endpoints, autoscaling, observability, function calling, structured outputs, and safety guardrails, allowing current customers to use it through model selection rather than new infrastructure, SDKs, billing arrangements, or vendor relationships.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.