Home / Companies / Zilliz / Blog / Post Details
Content Deep Dive

RAG Without OpenAI: BentoML, OctoAI and Milvus

Blog post from Zilliz

Post Details
Company
Date Published
Author
By Yujian Tang
Word Count
2,820
Company Posts That Month
21
Language
English
Hacker News Points
-
Post removed?
No
Summary

This tutorial demonstrates how to build retrieval augmented generation (RAG) applications using large language models (LLMs) without relying on OpenAI. The process involves serving embeddings with BentoML, inserting data into a vector database for RAG, setting up an LLM for RAG, and providing instructions to the LLM. Key components include BentoML for serving embeddings, OctoAI for accessing open-source models, and Milvus as the vector database. The example uses BentoML's Sentence Transformers Embeddings repository, a local Milvus instance using Docker Compose, and the Nous Hermes fine-tuned Mixtral model from OctoAI for RAG.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 51 2,722 279 102 +43%
RAG 24 1,867 232 78 +54%
LLM 15 3,669 412 154 +40%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.