Gemma-4 31B + vLLM on RTX 6000 PRO : A Real-Load Benchmark
Blog post from Hugging Face
Gemma-4 31B, a dense Transformer model developed by Google DeepMind, was evaluated on the vLLM serving engine using an RTX 6000 PRO GPU to benchmark its performance under varying concurrency levels from 12 to 24. The model, optimized for reasoning, coding, and multimodal understanding, was tested with a 4K-token context window to align with ShareGPT dataset requirements. Results showcased impressive throughput peaking at 1.17k tokens per second, with a median time to first token (TTFT) of approximately 0.7 seconds, although tail latency presented a challenge under heavy load, reaching up to 19 seconds for the p99 metric. Despite this, the server maintained low queue depths, indicating efficient handling of requests even at maximum concurrency, with long end-to-end latencies attributed primarily to the generation of lengthy outputs rather than server inefficiencies. The benchmarking was conducted using HexGrid.cloud, highlighting the platform's capability to deploy open models on dedicated GPUs effectively.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 2 | 5,601 | 1,340 | 262 | -2% |
| LLM | 1 | 6,196 | 1,155 | 243 | -32% |
| Vector Search | 1 | 1,895 | 382 | 133 | -16% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.