Open-sourcing evals for open-weight agents
Blog post from Cline
Cline argues that improving coding-agent performance on increasingly capable and affordable open-weight models requires disciplined evaluation rather than simply maximizing benchmark scores or reasoning tokens. Using the Harbor framework and Terminal-Bench’s 89 coding tasks, the company found its agent requests used 20–30% more tokens than efficient public harnesses despite strong scores, prompting a multi-objective approach balancing intelligence, cost, speed, and token efficiency. Its proposed “Hill Climber’s Checklist” recommends establishing a practical North Star metric, measuring score variance while eliminating infrastructure flakiness, isolating failures by task, model, and provider, testing whether greater reasoning effort actually improves outcomes, and maintaining private evaluation sets to reduce benchmark contamination and reward hacking. Examples show that a small number of tasks can account for most token consumption, identical harness changes can substantially help some model families while harming others, and provider routing alone can create large differences in success rates, token use, caching, and cost. The central conclusion is that there is no universally optimal agent harness or configuration; teams must repeatedly test, inspect traces, accept tradeoffs, and optimize for their own models, providers, and product goals.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Subagents | 2 | 276 | 88 | 41 | +39% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.