Home / Companies / Atlas Cloud / Blog / Post Details
Content Deep Dive

We Gave Qwen3.7-Plus 10 Real Bugs and 15 AIME Problems. It Outperformed the Flagship Model in Both.

Blog post from Atlas Cloud

Post Details
Company
Date Published
Author
Atlas Cloud
Word Count
3,433
Company Posts That Month
293
Language
English
Hacker News Points
-
Post removed?
No
Summary

An independent, single-afternoon evaluation compared Alibaba’s Qwen3.7-Plus, Qwen3.7-Max, and Qwen3.6-Plus across real-world bug repair, AIME 2025 math problems, speed, cost, and limited image understanding using a fixed Stirrup agent framework and external pytest verification. Qwen3.7-Plus completed all 10 BugFind repair tasks in the single run, while Max and 3.6-Plus completed nine, and it achieved 14 of 15 AIME problems with thinking enabled, matching Max’s score while responding substantially faster. The assessment found a roughly 3.55-fold throughput improvement over Qwen3.6-Plus, whose reasoning mode was too slow to finish the full math set, and argued that thinking should be enabled selectively because it improved hard-problem accuracy but added tokens and latency without benefiting simple questions. Qwen3.7-Plus was positioned as a lower-cost multimodal alternative to the text-focused Max, although an official image sample was incorrectly described despite success on a simple controlled image test. The authors emphasize that the findings are not a general model ranking because they come from one run, a small test set, one agent scaffold, limited vision testing, and no comparisons with competing model families, but recommend evaluating models on organizations’ own tasks with reproducible external checks.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 2 6,200 1,430 272 +10%
Real-time 2 6,055 1,444 270 -11%
LLM 1 6,292 1,205 252 -36%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.