Home / Companies / Neo4j / Blog / Post Details
Content Deep Dive

YouTube Transcripts Into Knowledge Graphs for RAG Applications

Blog post from Neo4j

Post Details
Company
Date Published
Author
Alex Gilmore
Word Count
1,770
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

This blog post explores how to scrape YouTube video transcripts into a knowledge graph for Retrieval Augmented Generation (RAG) applications. The project uses Google Cloud Platform, Neo4j, and LangChain to create a document from the transcript, store the resulting documents in a Neo4j graph database, and embed only the smaller child chunks of the text using SpaCy embeddings. The process involves setting up services such as Google Cloud Storage and Neo4j AuraDB instance, scraping transcripts from YouTube videos, chunking the transcripts into manageable pieces, loading the transcripts into the Neo4j graph database, and creating an index on the embedding property for vector search. The project demonstrates how to build a simple knowledge graph that can be used for RAG applications, with plans to explore building a basic RAG application in the next blog post.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 25 1,692 211 78 +87%
RAG 8 1,360 163 55 +97%
LLM 2 2,593 281 107 +38%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.