Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Driving model performance optimization: 2024 highlights

Blog post from Baseten

Post Details
Company
Date Published
Author
Pankaj Gupta
Word Count
1,530
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

In 2024, Baseten's Model Performance Team made significant breakthroughs in optimizing latency, scalability, quality, cost, functionality, and ease of use for high-volume real-world workloads. They adopted TensorRT-LLM as their core framework, leveraging its performance and incorporating features like Flash Attention, paged attention, and in-flight batching with SOTA CUDA kernels. The team also explored NVIDIA's Hopper architecture, particularly the H100 GPU, which offered exceptional performance thanks to large and high-bandwidth onboard memory, strong compute profiles, and excellent architectural features. Additionally, they implemented featureful inference servers, including guaranteed structured output, function calling, LoRA inference support, and innovations like Writing in the Margins for long-context retrieval accuracy. The team also developed automated tools like Engine Builder to streamline engine creation and deployment, reducing manual effort and improving efficiency. Notable achievements included optimizing custom LLMs, real-time AI phone calls, Whisper ASR, DeepSeek V3, and large model cold starts in under a minute. Looking ahead to 2025, the team is excited to broaden and deepen their work on speculative decoding, embeddings models, Blackwell GPU architecture, FP4 quantization, disaggregated serving, and more.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 20 3,709 434 145 +39%
Real-time 4 3,671 840 202 +19%
AI Model Fine-tuning 2 862 147 71 +81%
Vector Search 2 2,433 274 99 -40%
Local AI 1 17 11 8 +6%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.