GDP.xlsx: Can Agents Understand the Spreadsheets That Run the World?
Blog post from Surge AI
GDP.xlsx is a benchmark of 70 real-world Excel tasks across 12 professional domains designed to test whether AI agents can navigate and reason over complex workbooks rather than merely extract cell values. It emphasizes that spreadsheet logic is often distributed across tabs, formulas, lookup tables, formatting, comments, hidden rows, and embedded business rules, requiring models to determine which sources are authoritative and how information connects. The leading model cited, Gemini 4 Argon, achieved 38.3%, illustrating the difficulty of the benchmark. Case studies describe models overlooking decisive information on separate tabs, such as manufacturer-specific certification rules and lease-break clauses, even when they correctly identified related data elsewhere. Built with domain professionals from actual workplace workflows and reviewed against detailed grading rubrics, the benchmark aims to assess whether an AI can provide decisions that users could trust without extensive human guidance or auditing.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | 6 | No monthly metrics for this publish month. | |||
| AI Agents | 3 | 931 | 231 | 103 | -84% |
| LLM | 2 | 747 | 162 | 79 | -85% |
| Cost per task | 1 | 10 | 5 | 5 | -84% |
| Gemini 4 Argon | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.