Microsoft used Surge's human evaluation for MAI-Thinking-1
Blog post from Surge AI
Microsoft conducted a study using Surge's human evaluation services to assess the real-world effectiveness of their AI model, MAI-Thinking-1, as compared to Claude Sonnet 4.6. Instead of relying solely on benchmarks, which can sometimes be misleading, they employed blind human evaluations to determine which model users preferred across various tasks. This method highlighted the importance of human preference data, demonstrating that while benchmarks are essential for measuring specific capabilities, they do not fully capture the user experience. The study revealed that MAI-Thinking-1 was favored over its competitor in a significant number of tasks, emphasizing the value of human judgment in evaluating AI performance. Surge AI offers such evaluations to other developers to ensure their models not only perform well on paper but also provide a satisfactory user experience.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.