Your Agent Aced the Task. Will It Do It Again?
Blog post from Hugging Face
IBM researchers argue that conventional average accuracy metrics can obscure agent unreliability, citing a GPT-4.1 ReAct agent that achieved 77.4% Mean@5 accuracy on AppWorld but completed all five repeated runs for only 53.0% of tasks. They propose reporting Pass^k, which measures tasks consistently solved in every repetition, and define the difference between Mean@k and Pass^k as a consistency gap. To reduce this gap, the open-source ALTK-Evolve system adds a Consistency Analyzer that resamples individual decision points from a recorded trajectory to identify unstable, near-tied choices without requiring ground truth, model internals, live environment replay, or full task reruns. It then generates reusable consistency guidelines targeting vulnerable reasoning or tool-use decisions, such as validating search results or using robust counting methods. In tests on 168 AppWorld tasks, these guidelines raised Pass^5 from 53.0% to 69.0% and increased Mean@5 from 77.4% to 81.0%, reducing the consistency gap from 24.4 to 12 percentage points while also showing gains on related tasks.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.