Vals Web Search Index: Evaluating Search for Real Work
Blog post from Vals
Evaluating AI web-search tools is difficult because static benchmarks can be contaminated, may reward models’ pretrained knowledge rather than search use, and often contain narrow tasks unlike real professional work. The passage argues that widely used benchmarks such as BrowseComp and HLE have become less reliable as questions and answers leak online, while research suggests some models can answer many BrowseComp questions without tools and may use search chiefly to verify existing knowledge. It describes alternative efforts by search providers to use newer or dynamically generated tasks, then introduces the Vals Web Search Index as an independent evaluation intended to measure whether an agent can complete realistic knowledge-work tasks using a specific search tool. The index holds the model and evaluation harness constant while swapping search tools, scores final end-to-end answers rather than individual search-result rankings, and currently draws from expert-authored finance and legal benchmarks. Its private, continuously refreshed test set and zero-data-retention procedures are intended to limit memorization and leakage; reported no-search results were much lower than search-enabled performance. The authors contend that benchmarks should prioritize accurate completion of consequential real-world tasks over retrieval of isolated facts, since evaluation standards influence how search systems are optimized.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.