Home / Companies / PostHog / Blog / Post Details
Content Deep Dive

How we built automatic clustering for LLM traces

Blog post from PostHog

Post Details
Company
Date Published
Author
Andy Maguire
Word Count
3,195
Company Posts That Month
27
Language
-
Hacker News Points
-
Post removed?
No
Summary

In the evolving landscape of AI, the integration of traditional clickstream analytics with backend observability data has given rise to a new type of data, specifically focusing on LLM (Large Language Model) traces and generations. This data is invaluable for assessing AI agent performance and understanding user interactions, including emotional nuances. PostHog has developed a "Clustering" pipeline that processes these traces into actionable insights. The pipeline ingests raw traces, converts them into readable text, samples them, and uses LLMs for structured summarization, followed by embedding and clustering with techniques like UMAP and HDBSCAN. This process results in semantically rich clusters that are then labeled by an AI agent for meaningful interpretation. PostHog's solution balances accuracy and cost by employing sampling and structured outputs, enabling AI product teams to uncover usage patterns and potential issues efficiently. The clustering jobs feature allows for customizable clustering configurations, enhancing the ability to analyze specific data subsets. The system operates seamlessly within PostHog's AI Observability framework, offering both automatic and user-steered clustering capabilities.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 20 7,531 1,250 268 +26%
Vector Search 17 3,215 679 175 +33%
Observability 6 4,660 984 209 +14%
RAG 6 2,000 386 114 +12%
AI Agents 3 7,403 1,426 278 +69%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.