January 2020 Summaries
3 posts from Surge AI
Filter
Month:
Year:
Post Summaries
Back to Blog
HANDBOOK.md is a benchmark designed to evaluate the ability of AI agents to follow complex company handbooks in realistic environments, encompassing tasks across domains like Finance, Medical Billing, Insurance, Logistics, and HR. These tasks present AI agents with long, intricate policy documents that they must interpret and apply while navigating tools like email, Slack, Jira, and spreadsheets. Despite the assumption that AI can adhere to detailed instructions, top models currently succeed in less than 25% of these tasks, often failing to maintain rule fidelity, missing important authorization steps, and erroneously reporting compliance. The benchmark challenges models with unique, non-memorizable handbooks for each task, posing a significant test of their capacity to remain consistent and accurate over extended operations. This reflects a critical gap in current AI capabilities, highlighting the difficulty in ensuring models can reliably execute long, policy-driven tasks, which is crucial for their deployment in real-world enterprise scenarios.
Jan 01, 2020
3,421 words in the original blog post.
HANDBOOK.md is a benchmark designed to evaluate AI agents' ability to follow complex, realistic company handbooks in live enterprise environments, spanning domains such as Finance, Medical Billing, Insurance, Logistics, and HR. The benchmark challenges AI by requiring them to navigate tasks with extensive policy documents, which can range from 20 to 124 pages and come in various formats like PDF and Word. Despite these detailed instructions, current frontier AI models, including GPT-5.5 and Opus 4.8, struggle with these tasks, often failing to comply with crucial rules and achieving success rates below 25%. The benchmark highlights critical failure patterns, such as prioritizing immediate requests over standing policies and losing track of information over extended tasks. This reflects the ongoing challenge for AI models to adhere to complex, long-standing instructions in dynamic, real-world settings.
Jan 01, 2020
2,810 words in the original blog post.
Chartography is introduced as a novel benchmark designed to evaluate professional graphical reasoning skills, particularly focusing on whether advanced AI models can interpret specialized charts commonly used across fields like medicine, engineering, finance, and science. Unlike traditional benchmarks which focus on simple chart types such as bar or line graphs, Chartography includes complex formats like Sankey diagrams, candlestick charts, and contour maps, requiring models to perform tasks like estimating unlabeled values and interpreting complex geometries. Despite the capabilities of frontier models, which perform well on existing benchmarks, they struggle with Chartography, scoring below 50% due to challenges in visual estimation and domain-specific conventions. The benchmark emphasizes the importance of precise visual reasoning and expert grading tailored to the complexities of professional chart reading, highlighting the gap between current AI capabilities and the nuanced skills required for interpreting professional charts.
Jan 01, 2020
2,311 words in the original blog post.