Your tools work. Will the agent use them right?
Blog post from Webflow
Webflow developed an MCP evaluation harness to test how real AI agents use its tools for tasks such as building pages, managing CMS content, and publishing sites, addressing limitations of conventional tests that verify tool correctness but not agent decision-making. The harness runs plain-language, multi-step “stories” against disposable Webflow sites using agent hosts including Claude Code and OpenAI Codex CLI, then combines deterministic assertions, transcript-based LLM judging, visual evaluation, structural HTML signals, and Datadog-tracked artifacts to assess both outcomes and execution paths. Its findings show that tasks can pass despite poor tool experiences, including constraints requiring awkward workarounds, oversized tool responses, and silent CSS class renaming. Repeated runs also revealed that stable tool-call success can conceal substantial variation in design quality, while semantic HTML measurements supplement screenshot-based assessments of accessibility and maintainability. Testing across multiple agent hosts proved especially important after behavior that passed consistently with Claude failed with Codex until prompts explicitly prohibited unavailable Designer-canvas tools. The system now supports pre-release, nightly, and self-service evaluations, with planned improvements including coverage alerts, regression reporting, gold-reference quality benchmarks, and stories informed by aggregated real-world usage patterns.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.