Introducing Tuesday: A Frontier Index for AI at Work
Blog post from Surge AI
The Tuesday Index is a composite benchmark from Surge AI intended to measure how well frontier models handle ordinary professional work that requires multiple capabilities at once, such as interpreting charts, following long policies, using tools, reconciling constraints, exercising judgment, and communicating results clearly. It combines eight evaluations—Chartography, HANDBOOK.md, Antidote, Hemingway-bench, ComplexConstraints, GDP.pdf, CoreCraft, and Riemann-bench—to capture both foundational “floor” skills like instruction following and context retention and advanced “ceiling” reasoning capabilities. The index argues that current AI progress is uneven: models may solve difficult mathematics while still failing common workplace tasks involving documents, policies, charts, or messy organizational workflows. In its initial rankings, Fable 5 leads with a Tuesday Score of 66.8, narrowly ahead of GPT 5.6 Sol at 66.7, but the authors emphasize that no model has mastered a typical workday. The index is intended to expand as new benchmarks assess longer-term judgment, collaboration, delegation, artifact creation, prioritization, persuasion, and sustained usability.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 2 | 82 | 12 | 8 | - |
| LLM | 2 | 3,630 | 731 | 193 | -51% |
| AI Agents | 1 | 3,983 | 868 | 211 | -41% |
| MCP | 1 | 6,317 | 631 | 178 | -42% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.