Home / Companies / Neon / Blog / Post Details
Content Deep Dive

Comparing token economics across 42 AI models

Blog post from Neon

Post Details
Company
Date Published
Author
Carlota Soto
Word Count
2,249
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Neon compared 42 AI Gateway models on a synthetic workload of 100 customer-support tickets, measuring cost, token use, completion time, and whether responses met seven structured quality requirements involving JSON validity, classification, policy compliance, escalation, wording, and prohibited claims. GPT-5 Nano delivered the lowest cost per usable response and completed the workload in about three minutes, while GPT-5.3 Codex achieved the highest pass rate at 82%; Llama 3.1 8B Instruct was fastest at 1 minute 18 seconds but passed only 34 tickets, and GPT-5.5 Pro was the most expensive at roughly $8.34 and took 57 minutes. Results also showed that Qwen3.5 122B-A10B consumed the most tokens, while open-weight models had substantially lower median costs than proprietary models but somewhat lower median pass rates. The benchmark used isolated Neon branches for each model, a Python runner to submit and grade requests, Postgres for results, and a Next.js web application for reporting, with scheduled reruns planned. The authors emphasize that findings apply only to this workload and reflect real-world variability, but estimate that model selection could create an 86-fold difference in agent inference spending for organizations operating at large token volumes.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 5 No monthly metrics for this publish month.
Serverless 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.