Dollars per token considered harmful
Blog post from Modal
Self-hosted, open-source large language model inference should be evaluated primarily in dollars per request rather than dollars per token, according to the text, because application users pay for completed interactions rather than token consumption. While token-based pricing suits API providers whose costs scale with input and output size, teams operating their own models must instead connect infrastructure decisions to user workflows, latency expectations, concurrent request volumes, and business value. The text argues that estimating acceptable response times requires considering time to first token and generation speed for typical requests, while capacity planning depends on requests per second and the number of concurrent requests each model replica can handle. Since self-hosting costs are ultimately determined by compute expenses over time and the replicas needed to meet performance targets, token counts become an internal operational metric rather than the main pricing or strategic measure. Framing costs per request is presented as a clearer way to assess whether an LLM feature produces sufficient revenue, conversion, or customer value to justify its infrastructure expense.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 22 | 4,922 | 763 | 224 | +11% |
| Serverless | 1 | 1,048 | 263 | 99 | +36% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.