Home / Companies / Surge AI / Blog / February 2026

February 2026 Summaries

2 posts from Surge AI

Filter
Month: Year:
Post Summaries Back to Blog
EnterpriseBench, a suite of reinforcement learning environment benchmarks, evaluates AI agents on high-value job functions within realistic enterprise settings, using a startup called CoreCraft as the initial test environment. CoreCraft challenges agents with tasks such as navigating complex databases, managing customer interactions, and adhering to company policies, reflecting real-world enterprise operations. Despite the sophistication of state-of-the-art models like GPT-5.2 and Claude Opus 4.6, they solved fewer than 30% of the tasks, often faltering due to hallucinations and reasoning errors. Training improvements were seen with the GLM 4.6 model, which demonstrated gains in executing multi-step workflows and handling constraints. These advancements were not only evident within CoreCraft but also transferred to external benchmarks, suggesting the acquisition of generalizable skills. The initiative aims to expand by building environments for other job families, enhancing the practical applicability of AI in enterprise contexts.
Feb 19, 2026 4,038 words in the original blog post.
Hemingway-bench is a new AI writing leaderboard designed to enhance the evaluation of AI-generated writing by emphasizing genuine creativity, nuance, and depth over superficial metrics. Unlike traditional benchmarks such as EQ-Bench Creative Writing and LMArena, which often reward models for meeting basic structural criteria or favoring clickbait-style content, Hemingway-bench uses expert human evaluators to assess writing across a spectrum of real-world and frontier tasks. These tasks range from creative storytelling to business document writing, with models judged on dimensions such as creativity, coherence, and writing quality. The results highlighted Google's Gemini and Claude's Opus as top performers, with their strengths in creating engaging narratives and maintaining a human-like voice. The initiative aims to shift the focus from rewarding flashy, superficial prose to recognizing writing that offers depth, emotional resonance, and insightful commentary, thus redefining what constitutes quality in AI-generated content.
Feb 04, 2026 3,283 words in the original blog post.