Building a Golden Eval Dataset from Production Traffic
Blog post from OpenRouter
Golden eval datasets are curated, version-controlled collections of scrubbed production inputs paired with human-reviewed expected outputs or rubrics, designed to detect regressions from prompt, model, or provider changes before deployment. Unlike general benchmarks or synthetic-only tests, they reflect actual user traffic, product-specific failure modes, and changing usage patterns, though synthetic examples can supplement rare but important cases. The recommended workflow is to sample and anonymize production traffic, deduplicate and cluster examples for representative coverage, create clear binary and analytic grading criteria, validate the set against the current model to remove ambiguous items and repair flawed rubrics, then commit the dataset, rubric, and baseline results to Git and run evaluations in CI. Dataset size should prioritize coverage of meaningful failure modes, ranging from small diagnostic sets to hundreds or thousands of items for broad regression testing, with faster subsets for pull requests. The same frozen dataset can also compare candidate models through OpenRouter’s unified API, while accounting for differences in supported features, provider routing, cost, and context limits. LLM judges can automate rubric-based scoring but should be calibrated against human labels, use a different model than the one being tested, and not serve as the sole basis for high-stakes decisions.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 8 | 747 | 162 | 79 | -85% |
| Vector Search | 1 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.