Fable 5.1 Scores 68.7 on the Tuesday Work Index
Blog post from Surge AI
Surge AI’s comparison of recent frontier-model releases finds that Fable 5.1 leads overall on its Tuesday Work Index with a score of 68.7 and tops benchmarks for professional chart reasoning, policy-document instruction following, and enterprise agent performance, though its cost efficiency varies by task. Muse Spark 1.3 records the largest generation-to-generation index gain and leads ComplexConstraints, where it combines strong performance on interacting professional requirements with favorable costs, while showing uneven results on chart reasoning and advanced mathematics. Gemini 3.8 Flash reaches 61.1 on the index and is highlighted for cost-efficient hard reasoning, particularly after a 12-point improvement on Riemann-bench, although higher reasoning settings do not consistently improve its results across workloads. Together, the releases suggest that top-end capability continues to advance while lower-cost models narrow the gap, and GPT-6 Astra remains absent from the comparison pending completion of its evaluation.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 3 | 0 | 0 | 0 | -100% |
| LLM | 2 | 747 | 162 | 79 | -85% |
| AI Agents | 1 | 931 | 231 | 103 | -84% |
| Cost per task | 1 | 10 | 5 | 5 | -84% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.