NEEDLE: The benchmark your search engine can't memorize
Blog post from Keenable
NEEDLE is presented as a continuously refreshed, open-source benchmark for evaluating web search engines used by AI agents, designed to avoid the leakage, memorization, and overfitting problems associated with static benchmarks. Its queries span breaking news, finance, scholarly research, rare long-tail entities, and legal materials, drawing on live public data sources and agent search logs, with daily or hourly updates and publicly available code, tasks, and metrics. The benchmark compares Keenable with Google, Bing, Brave, Tavily, Parallel, and Exa, while also estimating an upper bound based on the best results collectively returned by all engines. The authors argue that rare, natural-language, and multi-step research queries remain particularly difficult and are most representative of real agent traffic. They further contend that independently owned search indexes are essential for improving retrieval quality, reducing latency, and providing unique coverage, whereas engines built around human engagement signals such as clicks, videos, and simplified queries may be poorly suited to agents, which search iteratively, use advanced operators, and can efficiently inspect authoritative sources.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.