Home / Companies / Surge AI / Blog / Post Details
Content Deep Dive

Qwen 3.8 Max Scores 58.7 on the Tuesday Work Index

Blog post from Surge AI

Post Details
Company
Date Published
Author
-
Word Count
1,250
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Qwen 3.8 Max achieved a score of 58.7 on the Tuesday Work Index, improving 8.6 points over Qwen 3.7 Max and 22.4 points over Qwen 3.5 Plus, placing it near several frontier-model operating points but below the highest-scoring models, Fable 5 Adaptive Max and GPT 5.6 Sol Max. Its largest gains appeared in structured professional instruction following and graphical reasoning: it scored 45.5% on the ComplexConstraints benchmark, reaching roughly 90% of GPT 5.6 Sol Max’s leading score at 32% of its reported evaluation cost, and improved to 29.1% on Chartography after relatively flat earlier results. Performance was uneven across benchmarks, however, with no improvement on Riemann-bench research mathematics and a decline in Antidote’s expert-graded answer-quality Elo score. Overall, the evaluation characterizes Qwen 3.8 Max as the strongest measured Qwen model so far, with particularly competitive cost-performance on some professional workloads despite remaining behind the top frontier systems in absolute capability.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Gemini 3.7 Flash 2 82 12 8 -
LLM 2 3,630 731 193 -51%
AI Agents 1 3,983 868 211 -41%
MCP 1 6,317 631 178 -42%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.