Cost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models
Blog post from Arize
A benchmarking study by Arize and Fireworks evaluated 10 AI models, including both open and closed types, using 40 real agent tasks over 2,400 runs to determine the cost per successful task, rather than relying on token pricing. This approach revealed that the price per successful task, which includes costs from retries and failures, is a more reliable metric for evaluating model efficiency and productivity. The study found that models with lower cost per successful task, like gpt-oss-120b, were more cost-effective despite lower pass rates compared to models like GPT-5.5 and Kimi K3, which performed better on difficult tasks but at a higher cost per success. The results emphasized the importance of routing tasks based on difficulty, suggesting that using cost-effective models for simpler tasks and reserving more capable, expensive models for complex tasks can optimize both cost and performance. The study highlighted that the label of a model being open or closed was not predictive of its performance or cost-effectiveness; instead, the choice should be based on capability and the specific requirements of the task at hand.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.