November 2025 Summaries
2 posts from Surge AI
Filter
Month:
Year:
Post Summaries
Back to Blog
In 2025, the focus of artificial intelligence development shifted towards creating agents capable of performing economically valuable tasks in realistic environments, moving beyond simple chat interfaces. Despite advancements, models like GPT-5 and Claude Sonnet 4.5 still struggled with over 40% of tasks in reinforcement learning (RL) environments, highlighting the challenges in developing generally intelligent agents. These environments, exemplified by Corecraft, Inc., are designed to mimic real-world tasks, such as customer support, requiring models to perform complex operations like tool use, goal formation, and adaptability. A hierarchy of agentic capabilities was identified, ranging from basic tool use to common-sense reasoning, with current models exhibiting varied levels of proficiency. While newer models like GPT-5.2 and Claude Opus 4.5 showed modest improvements, they still face significant obstacles in achieving human-level common-sense reasoning, which remains a critical barrier to their real-world applicability. The year marked significant progress in AI agents' reliability and coherence, setting the stage for further exploration into their potential to match human intelligence, although the timeline for closing this gap remains uncertain.
Nov 03, 2025
4,073 words in the original blog post.
In a comprehensive evaluation of advanced language models in the finance domain, experts tested three models—GPT-5, Claude Sonnet 4.5, and Gemini 2.5 Pro—across over 200 scenarios, revealing both their potential and limitations. While GPT-5 emerged as the most effective, excelling in 47% of tasks and outperforming the others in head-to-head comparisons, all models demonstrated significant shortcomings. These included failure to account for real-world financial constraints, poor multi-step workflow execution, and errors in file handling and domain calibration. Specific examples, such as creating PowerPoint presentations for market crash scenarios and updating financial forecasts in Excel, highlighted these deficiencies, with GPT-5 being the only model to produce a nearly complete deliverable, yet still lacking in areas like risk mitigation commentary. The study underscored the models' sophistication but also their systematic gaps, emphasizing the need for high-quality, real-world training data to bridge the gap between theoretical knowledge and practical financial expertise.
Nov 03, 2025
3,212 words in the original blog post.