Home / Companies / Atlas Cloud / Blog / Post Details
Content Deep Dive

DeepSeek Harness vs Hermes: Which One Burns More Tokens?

Blog post from Atlas Cloud

Post Details
Company
Date Published
Author
Atlas Cloud
Word Count
3,946
Company Posts That Month
133
Language
English
Hacker News Points
-
Post removed?
No
Summary

A controlled comparison of DeepSeek Harness (dsh) and Hermes Agent using the same DeepSeek V4 Pro model, API endpoint, key, prompt, machine, and empty working directories found that both successfully created and tested a single-file Breakout game, but dsh completed the task much faster and with far fewer tokens. Dsh took about 121 seconds, eight logged tool calls, and 132,600 prompt tokens, while Hermes took about 780 seconds, roughly 35–38 calls, and 1.11 million prompt tokens, producing an estimated 8.4-fold prompt-token and cost difference under the stated pricing method. The account attributes much of the gap to harness behavior, including different system-prompt and tool-schema sizes, repeated conversation-context transmission, step counts, retries, and context-window settings; even a prompt asking only for “OK” used 10,898 tokens in dsh and 13,892 in Hermes. It provides configuration and measurement guidance for reproducing the test with a shared OpenAI-compatible endpoint, emphasizing explicit context settings, compatible tool support, machine-readable usage logs, and provider billing as the ultimate source of truth. The comparison also notes that dsh is positioned as a leaner coding-oriented developer preview with session replay but no persistent memory or messaging integrations, whereas Hermes offers long-term memory, reusable skills, scheduling, chat channels, and dashboards at substantially higher measured token use, suggesting that some users may combine Hermes for persistent coordination with dsh for coding tasks.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.