Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How we built high-throughput embedding, reranker, and classifier inference with TensorRT-LLM

Blog post from Baseten

Post Details
Company
Date Published
Author
Michael Feil, Philip Kiely
Word Count
2,035
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

We built Baseten Embedding Inference (BEI), an optimized inference runtime leveraging TensorRT-LLM to significantly boost throughput and minimize latency for embedding, reranker, and classification models. BEI outperforms the competition by a large margin, offering double the throughput of previous industry standards in batch inference and improved latency for real-time queries. The core engine for BEI is TensorRT-LLM, which offers exceptional performance and consistent throughput without the risk of OOM errors. BEI has four main components: the model server, tokenizer, batch manager, and TensorRT-LLM inference engine. The runtime benefits from optimized inference engines, such as XQA kernel and layer fusing, as well as quantization to FP8, which provides a 50% or more gain in throughput while retaining >99% cosine similarity to outputs from non-quantized models. BEI supports traffic-based autoscaling, deployment on multiple clouds and regions, and reduced communication overhead between models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 44 1,879 278 111 +3%
LLM 22 4,855 541 180 +51%
Real-time 5 4,629 997 226 +44%
RAG 4 1,499 228 73 +7%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.