Home / Companies / Voxel51 / Blog / Post Details
Content Deep Dive

AI Agent Cost Compared: Claude, GPT-5.1, and Gemini 3.1

Blog post from Voxel51

Post Details
Company
Date Published
Author
-
Word Count
2,191
Company Posts That Month
18
Language
English
Hacker News Points
-
Post removed?
No
Summary

Voxel51 compared Claude Sonnet 5, Claude Opus 5, GPT-5.1, and Gemini 3.1 Pro by running identical FiftyOne Agent conversations with the same prompts, tools, skills, and dataset, using provider-reported API token usage to measure actual costs. Across three core tasks involving image description, dataset summarization, and creating a dataset filter, GPT-5.1 was the least expensive at $0.109 and Claude Opus 5 the most expensive at $1.700, while Claude Sonnet 5 achieved the highest quality score of 8.67 out of 10 for $0.331. All models passed the task rubric, but their costs, token consumption, and response quality differed substantially, showing that token pricing alone does not predict operational expense. More demanding tasks highlighted failure behavior as a major cost risk: GPT-5.1 generally stopped quickly when unsuccessful, whereas Sonnet and Gemini sometimes repeatedly pursued dead ends, producing single-task costs above $7. The results are limited by single runs, an isolated sandbox dataset, conservative Claude cost estimates due to cache accounting, and a headless test setup that affected some browser-dependent operations, but they suggest GPT-5.1 is the value-oriented choice and Claude Sonnet 5 offers the strongest quality-per-dollar balance.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 11 931 231 103 -84%
LLM 5 747 162 79 -85%
AI Guardrails 4 35 22 12 -94%
MCP 3 2,241 148 72 -74%
Cost per task 2 10 5 5 -84%
Vector Search 2 265 57 33 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.