Home / Companies / CircleCI / Blog / Post Details
Content Deep Dive

Deploying a multimodal RAG application with Gemma 3 and CircleCI on GKE

Blog post from CircleCI

Post Details
Company
Date Published
Author
Armstrong Asenavi
Word Count
5,207
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Retrieval-Augmented Generation (RAG) is a transformative approach for enhancing interactions with Large Language Models (LLMs) by grounding responses in external knowledge to improve accuracy and reduce errors. Traditional RAG systems are limited to text processing, but multimodal RAG overcomes this by integrating text, images, and potentially audio, creating a more comprehensive understanding similar to human sensory integration. This tutorial guides the construction of a multimodal RAG application using Google’s Gemma 3 model served via Ollama, which processes PDF documents containing both text and images. The application employs Qdrant as a vector store, creates an interactive UI with Streamlit, and is deployed to Google Kubernetes Engine (GKE) using CircleCI, resulting in a scalable system that allows users to query document content irrespective of format. The process involves setting up a GKE cluster, creating Docker containers, and deploying services using Kubernetes, while CircleCI automates the build and deployment pipeline to ensure a streamlined workflow.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 40 1,504 310 125 -10%
RAG 37 1,006 206 82 -15%
Kubernetes 25 893 168 80 -9%
LLM 3 3,636 538 190 -7%
Serverless 2 842 169 80 +38%
Real-time 1 4,065 968 231 -6%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.