DeepSeek wins IMO Gold on 12 cents
Blog post from Cline
An evaluation using Cline tested eight AI models on six IMO 2026 problems, with anonymized proofs graded on the 0–7 IMO scale by two independent model judges and a third used to resolve disagreements. GPT-5.6 Sol achieved a perfect 42/42 score, while Claude Fable 5 scored 41/42, but the reported lowest-cost gold-level result came from DeepSeek V4 Flash, which earned 30/42 against a 29-point gold cutoff for $0.1215, compared with $17.1956 for Claude Fable 5’s best run. DeepSeek V4 Pro and MiMo V2.5 Pro also reached 30/42, while the stated human median was 16/42. The authors emphasize that harness bugs were corrected, internet access was disabled, solution traces were reviewed, and the new problems were not expected to be in training data, while acknowledging that model scores can vary with prompting, retries, and provider reliability and that reported prices reflect only the best runs. They present the results as evidence of rapidly improving mathematical reasoning, particularly among open-weight models, and contrast them with earlier milestones such as AlphaProof’s 2024 silver-level performance and Gemini Deep Think’s reported gold-level result in 2025.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.