How Similarweb Evaluates Long-Form Agent Research Reports with LangSmith
Blog post from LangChain
Liora Korni, a Senior AI Engineer at SimilarWeb, discusses the challenges and methodologies for evaluating agentic systems like Similarweb Data Studio. Unlike traditional software where outputs are consistent, agentic systems produce varied results from the same inputs due to their dynamic nature. Similarweb Data Studio uses LangSmith to evaluate these systems, differentiating between deterministic checks and LLM-as-a-judge scoring to assess outputs. The deterministic checks ensure the correct tools are used, while the LLM-as-a-judge scoring assesses the quality and meaning of the outputs through rubric prompts and feedback. Korni emphasizes that evaluation must be integrated into the product architecture to understand quality changes, avoid miscalibration, and properly attribute scores to specific criteria. Misaligned rubrics can lead to incorrect judgments, as experienced with their Deep Research evaluation, highlighting the importance of aligning evaluations with desired behaviors. By combining golden answers, rubrics, faithfulness checks, and A/B comparisons, SimilarWeb creates a comprehensive evaluation framework that informs product decisions and optimizes agentic system performance.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 7,115 | 1,261 | 236 | +13% |
| Harness engineering | 1 | 222 | 129 | 60 | -13% |
| Observability | 1 | 3,826 | 727 | 190 | -10% |
| RAG | 1 | 1,170 | 274 | 98 | +16% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.