MLPerf Inference v6.1: pioneering agent, VLM benchmarks
Blog post from Lambda
Lambda reports MLPerf Inference v6.1 results highlighting advances in large-language-model, vision-language-model, and agentic inference on NVIDIA Blackwell-based systems. Its four-NVIDIA Blackwell Ultra GPU configuration achieved 65,511 tokens per second offline and 58,195 tokens per second in the server scenario for GPT-OSS 120B, leading comparable offline submissions and improving throughput by roughly 8.8% over its v6.0 results on identical hardware through software optimizations. Lambda also submitted Qwen3-VL-235B-A22B-Instruct, its first vision-language-model benchmark, with an HGX B200 system reaching 101.56 offline queries per second and 69.41 server queries per second. In the open division, it used an eight-GPU HGX B200 system to run Kimi K2.6, a mixture-of-experts model with more than one trillion parameters, in MLPerf’s new agentic workload, completing all 1,007 replay turns with 86.83% BFCL v4 accuracy and a mean per-turn latency of 770.8 milliseconds. The company attributes performance gains to both Blackwell Ultra hardware capabilities and optimizations including TensorRT enhancements, CUDA graphs, model quantization, KV-cache tuning, and automated expert-kernel selection, while noting that all results remain subject to final MLCommons review.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Serverless | 12 | 156 | 54 | 28 | -80% |
| LLM | 4 | 747 | 162 | 79 | -85% |
| AI Model Fine-tuning | 1 | 139 | 28 | 14 | -75% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.