How we built automatic clustering for LLM traces
Blog post from PostHog
In the evolving landscape of AI, the integration of traditional clickstream analytics with backend observability data has given rise to a new type of data, specifically focusing on LLM (Large Language Model) traces and generations. This data is invaluable for assessing AI agent performance and understanding user interactions, including emotional nuances. PostHog has developed a "Clustering" pipeline that processes these traces into actionable insights. The pipeline ingests raw traces, converts them into readable text, samples them, and uses LLMs for structured summarization, followed by embedding and clustering with techniques like UMAP and HDBSCAN. This process results in semantically rich clusters that are then labeled by an AI agent for meaningful interpretation. PostHog's solution balances accuracy and cost by employing sampling and structured outputs, enabling AI product teams to uncover usage patterns and potential issues efficiently. The clustering jobs feature allows for customizable clustering configurations, enhancing the ability to analyze specific data subsets. The system operates seamlessly within PostHog's AI Observability framework, offering both automatic and user-steered clustering capabilities.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 20 | 7,531 | 1,250 | 268 | +26% |
| Vector Search | 17 | 3,215 | 679 | 175 | +33% |
| Observability | 6 | 4,660 | 984 | 209 | +14% |
| RAG | 6 | 2,000 | 386 | 114 | +12% |
| AI Agents | 3 | 7,403 | 1,426 | 278 | +69% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.