Home / Companies / Vals / Blog / Post Details
Content Deep Dive

Vals Web Search Index: Evaluating Search for Real Work

Blog post from Vals

Post Details
Company
Date Published
Author
Nikil Ravi
Word Count
1,712
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Evaluating AI web-search tools is difficult because static benchmarks can be contaminated, may reward models’ pretrained knowledge rather than search use, and often contain narrow tasks unlike real professional work. The passage argues that widely used benchmarks such as BrowseComp and HLE have become less reliable as questions and answers leak online, while research suggests some models can answer many BrowseComp questions without tools and may use search chiefly to verify existing knowledge. It describes alternative efforts by search providers to use newer or dynamically generated tasks, then introduces the Vals Web Search Index as an independent evaluation intended to measure whether an agent can complete realistic knowledge-work tasks using a specific search tool. The index holds the model and evaluation harness constant while swapping search tools, scores final end-to-end answers rather than individual search-result rankings, and currently draws from expert-authored finance and legal benchmarks. Its private, continuously refreshed test set and zero-data-retention procedures are intended to limit memorization and leakage; reported no-search results were much lower than search-enabled performance. The authors contend that benchmarks should prioritize accurate completion of consequential real-world tasks over retrieval of isolated facts, since evaluation standards influence how search systems are optimized.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.