Choosing an AI model: one prompt, 11 models, very different results
Blog post from Netlify
Netlify’s partnership with OpenRouter expands its AI Gateway to support OpenRouter’s model catalog and adds open models including Kimi K3, GLM 5.2, and DeepSeek V4 to Agent Runners through the OpenCode coding agent, alongside existing Claude, OpenAI, and Gemini options. To compare model performance, Netlify used its open-source AXIS evaluation framework and ran identical prompts three times across models, assessing generated websites for functional correctness, appropriate use of platform features, and credit consumption. In an initial test to build a static neighborhood coffee-shop page, Claude Opus 5 generally produced the most detailed designs but had highly variable and sometimes very high costs, while models such as GPT 5.6 Terra, GLM 5.2, and DeepSeek V4 Flash offered lower-cost alternatives with differing levels of design quality and occasional implementation issues. The results suggest that model choice depends on whether users prioritize polished, self-directed output or a less expensive iterative workflow, while future comparisons will examine more complex applications involving databases, uploads, AI features, authentication, validation, and code quality.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 3,630 | 731 | 193 | -51% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.