Hill climbing to glory: using evals to improve AI error rate by 7x
Blog post from Hex
Hex reports using an evaluation-driven “hill-climbing” process to develop Quick Edits, a feature that uses smaller AI models to make simple chart changes or hand complex requests to its primary agent. Starting with real user requests and expanding to 1,800 varied cases with chart-state fixtures, the team prioritized reducing incorrect edits over unnecessary handoffs because unintended chart changes pose a greater trust risk. Automated agents tested proposed changes against hidden holdout cases, reducing the wrong-edit rate from 21% to 3%, while repeated runs helped control for model-output variability. More than half of the improvements came from strengthening the output validator to repair nearly correct responses and enforce chart constraints, rather than from prompt modifications; clearer property labels and moving certain logic out of the model also improved results. The company tested across models and providers, found that 10 attempts per case provided stable measurements, and kept evaluation runs inexpensive and fast enough to support continual changes to prompts, validators, chart properties, handoff rules, and user experience.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| GPT-6 Astra | 1 | No monthly metrics for this publish month. | |||
| LLM | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.