How to analyze AI agent usage patterns to build eval datasets (2026)
Blog post from Braintrust
AI agents generate vast amounts of production traces, but creating effective evaluation datasets from these logs is challenging due to the randomness of sampling and the limitations of hand-written tests. By analyzing production usage patterns, particularly focusing on tasks, user sentiment, and issues, product and engineering teams can identify significant workflows, user frustrations, and recurrent failures to construct more representative evaluation datasets. Braintrust Topics facilitates this process by clustering agent traces based on task intent, sentiment, and issue type, enabling teams to convert high-value trace groups into evaluation datasets that mirror real-world agent interaction. This approach ensures better coverage of potential production risks, unlike random sampling which may overlook rare but impactful failure modes. The system allows for continuous updates and alignment with current production behavior by regenerating patterns daily, and it supports the use of custom facets to address domain-specific questions. The ultimate goal is to create evaluation suites that are manageable in size but comprehensive in risk coverage, thereby improving the reliability and user experience of AI agents.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Agents | 4 | 6,119 | 1,396 | 266 | +24% |
| Observability | 1 | 4,230 | 776 | 198 | +24% |
| Vector Search | 1 | 1,897 | 384 | 134 | -16% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.