Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How we built BEI: high-throughput embedding, reranker, and classifier inference

Blog post from Baseten

Post Details
Company
Date Published
Author
Amir Haghighat 4 others
Word Count
2,111
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Baseten Embedding Inference (BEI) is an optimized runtime developed to enhance throughput and reduce latency for embedding, reranking, and classification models using TensorRT-LLM. The solution aims to address challenges posed by the increasing size of modern embedding models, which have evolved from BERT-based architectures to larger LLM-based models. BEI achieves significant performance improvements, offering up to 2.05 times the throughput of existing solutions while maintaining low latency and high concurrency for real-time queries. It supports various architectures, including newer models, and uses techniques like batching, sequence packing, and FP8 quantization to optimize performance further. Additionally, BEI incorporates infrastructure enhancements, such as traffic-based autoscaling and asynchronous inference, to handle high-throughput workloads efficiently. These developments position BEI as a leading option for deploying low-latency, high-throughput embedding models, offering developers flexibility and improved system performance.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 47 1,836 305 108 +20%
LLM 20 4,152 612 181 +19%
Real-time 5 4,668 1,055 221 +15%
RAG 4 984 209 73 -16%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.