Qwen 3.8 Max Scores 58.7 on the Tuesday Work Index
Blog post from Surge AI
Qwen 3.8 Max achieved a score of 58.7 on the Tuesday Work Index, improving 8.6 points over Qwen 3.7 Max and 22.4 points over Qwen 3.5 Plus, placing it near several frontier-model operating points but below the highest-scoring models, Fable 5 Adaptive Max and GPT 5.6 Sol Max. Its largest gains appeared in structured professional instruction following and graphical reasoning: it scored 45.5% on the ComplexConstraints benchmark, reaching roughly 90% of GPT 5.6 Sol Max’s leading score at 32% of its reported evaluation cost, and improved to 29.1% on Chartography after relatively flat earlier results. Performance was uneven across benchmarks, however, with no improvement on Riemann-bench research mathematics and a decline in Antidote’s expert-graded answer-quality Elo score. Overall, the evaluation characterizes Qwen 3.8 Max as the strongest measured Qwen model so far, with particularly competitive cost-performance on some professional workloads despite remaining behind the top frontier systems in absolute capability.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 2 | 82 | 12 | 8 | - |
| LLM | 2 | 3,630 | 731 | 193 | -51% |
| AI Agents | 1 | 3,983 | 868 | 211 | -41% |
| MCP | 1 | 6,317 | 631 | 178 | -42% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.