We Gave Qwen3.7-Plus 10 Real Bugs and 15 AIME Problems. It Outperformed the Flagship Model in Both.
Blog post from Atlas Cloud
In 2026, Alibaba's Qwen3.7-Max and Qwen3.7-Plus models were released, with Qwen3.7-Plus being positioned as a cost-effective multimodal model and Qwen3.7-Max as the text flagship. Available on Alibaba Cloud, these models underwent rigorous testing to evaluate their performance in tasks such as automatic bug repair, math problem-solving, and multimodality. Qwen3.7-Plus demonstrated notable improvements over its predecessor, Qwen3.6-Plus, with a 3.55x increase in throughput and lower latency in math tasks when compared to Qwen3.7-Max, although it struggled with complex visual tasks. Despite these advances, the evaluation highlighted the importance of careful task-specific model selection and the potential cost savings of dynamically enabling the models' "thinking" mode based on task difficulty. This assessment underscores the need for reproducible testing and evidence-based decision-making in adopting AI models for production environments.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.