Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Gemma-4 31B + vLLM on RTX 6000 PRO : A Real-Load Benchmark

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Nikhil K.
Word Count
786
Company Posts That Month
94
Language
-
Hacker News Points
-
Post removed?
No
Summary

Gemma-4 31B, a dense Transformer model developed by Google DeepMind, was evaluated on the vLLM serving engine using an RTX 6000 PRO GPU to benchmark its performance under varying concurrency levels from 12 to 24. The model, optimized for reasoning, coding, and multimodal understanding, was tested with a 4K-token context window to align with ShareGPT dataset requirements. Results showcased impressive throughput peaking at 1.17k tokens per second, with a median time to first token (TTFT) of approximately 0.7 seconds, although tail latency presented a challenge under heavy load, reaching up to 19 seconds for the p99 metric. Despite this, the server maintained low queue depths, indicating efficient handling of requests even at maximum concurrency, with long end-to-end latencies attributed primarily to the generation of lengthy outputs rather than server inefficiencies. The benchmarking was conducted using HexGrid.cloud, highlighting the platform's capability to deploy open models on dedicated GPUs effectively.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 2 5,601 1,340 262 -2%
LLM 1 6,196 1,155 243 -32%
Vector Search 1 1,895 382 133 -16%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.