Home / Companies / Surge AI / Blog / August 2026

August 2026 Summaries

4 posts from Surge AI

Filter
Month: Year:
Post Summaries Back to Blog
DeepSeek V4 Pro achieved 59.7 on Surge AI’s Tuesday Work Index, improving substantially over its preview versions and surpassing several competing models, although Fable 5 and GPT 5.6 Sol Max remain ahead. Across benchmarks measuring professional instruction following, enterprise agents, research mathematics, and writing, it was presented as competitive with frontier systems, with particular emphasis on cost efficiency. On ComplexConstraints, it scored 42.1%, or roughly 83% of the leading score, for about $34 per run—under 10% of the leading system’s cost—while on Riemann-bench it matched Grok 4.5’s 38.4% score for $11.04 compared with $122.38, placing it on the cost-performance Pareto frontier. The assessment characterizes DeepSeek V4 Pro as a strong open-weight option for workloads requiring complex reasoning at lower cost, while noting that it does not match the highest absolute scores of leading frontier models.
Aug 27, 2026 1,355 words in the original blog post.
Qwen 3.8 Max achieved a score of 58.7 on the Tuesday Work Index, improving 8.6 points over Qwen 3.7 Max and 22.4 points over Qwen 3.5 Plus, placing it near several frontier-model operating points but below the highest-scoring models, Fable 5 Adaptive Max and GPT 5.6 Sol Max. Its largest gains appeared in structured professional instruction following and graphical reasoning: it scored 45.5% on the ComplexConstraints benchmark, reaching roughly 90% of GPT 5.6 Sol Max’s leading score at 32% of its reported evaluation cost, and improved to 29.1% on Chartography after relatively flat earlier results. Performance was uneven across benchmarks, however, with no improvement on Riemann-bench research mathematics and a decline in Antidote’s expert-graded answer-quality Elo score. Overall, the evaluation characterizes Qwen 3.8 Max as the strongest measured Qwen model so far, with particularly competitive cost-performance on some professional workloads despite remaining behind the top frontier systems in absolute capability.
Aug 20, 2026 1,250 words in the original blog post.
The Tuesday Index is a composite benchmark from Surge AI intended to measure how well frontier models handle ordinary professional work that requires multiple capabilities at once, such as interpreting charts, following long policies, using tools, reconciling constraints, exercising judgment, and communicating results clearly. It combines eight evaluations—Chartography, HANDBOOK.md, Antidote, Hemingway-bench, ComplexConstraints, GDP.pdf, CoreCraft, and Riemann-bench—to capture both foundational “floor” skills like instruction following and context retention and advanced “ceiling” reasoning capabilities. The index argues that current AI progress is uneven: models may solve difficult mathematics while still failing common workplace tasks involving documents, policies, charts, or messy organizational workflows. In its initial rankings, Fable 5 leads with a Tuesday Score of 66.8, narrowly ahead of GPT 5.6 Sol at 66.7, but the authors emphasize that no model has mastered a typical workday. The index is intended to expand as new benchmarks assess longer-term judgment, collaboration, delegation, artifact creation, prioritization, persuasion, and sustained usability.
Aug 18, 2026 1,855 words in the original blog post.
The research focused on training the Qwen3.5-122B-A10B model on Long-Horizon Multi-Tool Agent Tasks, which are environments designed for complex office work involving documents, spreadsheets, and planning, but not coding. Despite this, the model improved its coding capabilities, highlighting an unexpected transfer of skills. The training emphasized goal-directed execution, involving defining goals, selecting actions, observing results, and updating the working state, which enhanced the model's ability to manage tasks with layered and interdependent goals. This process mirrors how humans tackle complex projects, such as planning events, where managing dependencies and maintaining overarching goals are crucial. The study suggests that such training can teach general capabilities beyond the specific domain, as the model showed improved performance on software-engineering tasks it had not specifically trained for, demonstrating the utility of well-designed datasets in fostering broad skill acquisition.
Aug 03, 2026 2,948 words in the original blog post.