Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

The Baseten Inference Stack at NVIDIA Dynamo Day

Blog post from Baseten

Post Details
Company
Date Published
Author
Rachel Rapp
Word Count
1,098
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

The Baseten Inference Stack utilizes a combination of open-source tools like NVIDIA Dynamo and in-house innovations to optimize generative AI workloads with the lowest latency and highest throughput. By leveraging Dynamo, which is framework-agnostic and regularly updated, Baseten can integrate various inference engines tailored to specific models and use cases. The use of Dynamo facilitates improvements in system-level inference performance through optimizations like disaggregated serving, KV cache-aware routing, and KV cache offloading, leading to significant reductions in latency and increases in throughput. Baseten's engineers contribute to the Dynamo ecosystem by offering enhancements and new features, which were highlighted during NVIDIA's Dynamo Day event. These strategies enable Baseten to achieve a 99.99% reliability rate in their AI model performance, while also supporting multimodal model serving by extending the capabilities of Dynamo in handling complex AI workloads.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 11 5,138 781 181 +34%
Vector Search 1 2,212 422 133 +33%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.