How we built high-throughput embedding, reranker, and classifier inference with TensorRT-LLM
Blog post from Baseten
We built Baseten Embedding Inference (BEI), an optimized inference runtime leveraging TensorRT-LLM to significantly boost throughput and minimize latency for embedding, reranker, and classification models. BEI outperforms the competition by a large margin, offering double the throughput of previous industry standards in batch inference and improved latency for real-time queries. The core engine for BEI is TensorRT-LLM, which offers exceptional performance and consistent throughput without the risk of OOM errors. BEI has four main components: the model server, tokenizer, batch manager, and TensorRT-LLM inference engine. The runtime benefits from optimized inference engines, such as XQA kernel and layer fusing, as well as quantization to FP8, which provides a 50% or more gain in throughput while retaining >99% cosine similarity to outputs from non-quantized models. BEI supports traffic-based autoscaling, deployment on multiple clouds and regions, and reduced communication overhead between models.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 44 | 2,157 | 323 | 132 | +11% |
| LLM | 22 | 5,694 | 663 | 215 | +42% |
| Real-time | 5 | 5,174 | 1,177 | 267 | +34% |
| RAG | 4 | 1,706 | 255 | 85 | +12% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.