Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How we run GPT OSS 120B at 500+ tokens per second on NVIDIA GPUs

Blog post from Baseten

Post Details
Company
Date Published
Author
Amir Haghighat 4 others
Word Count
938
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

Achieving state-of-the-art (SOTA) latency and throughput for the GPT OSS 120B model on NVIDIA GPUs involves a complex process of performance optimization, including experimentation, bug fixing, and benchmarking. The Baseten Inference Stack significantly contributes to this endeavor by allowing rapid performance improvements through its flexible architecture and the expertise of its model performance engineering team. Upon the model's release, engineers work in parallel using different inference frameworks such as TensorRT-LLM, vLLM, and SGLang, ensuring compatibility with Hopper and Blackwell GPU architectures. The team addresses compatibility bugs and optimizes model configurations, choosing Tensor Parallelism for better latency and leveraging TensorRT-LLM MoE Backend for enhanced performance. These efforts lead to significant improvements, including adding 100 tokens per second while maintaining 100% uptime, and highlight the importance of inference optimization for immediate improvements in latency and throughput. The team continues to explore new methods like speculative decoding to further enhance model performance, with a focus on providing efficient solutions for developers looking to optimize their models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 10 3,922 600 189 -6%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.