DeepSeek Harness vs Hermes: Which One Burns More Tokens?
Blog post from Atlas Cloud
A controlled comparison of DeepSeek Harness (dsh) and Hermes Agent using the same DeepSeek V4 Pro model, API endpoint, key, prompt, machine, and empty working directories found that both successfully created and tested a single-file Breakout game, but dsh completed the task much faster and with far fewer tokens. Dsh took about 121 seconds, eight logged tool calls, and 132,600 prompt tokens, while Hermes took about 780 seconds, roughly 35–38 calls, and 1.11 million prompt tokens, producing an estimated 8.4-fold prompt-token and cost difference under the stated pricing method. The account attributes much of the gap to harness behavior, including different system-prompt and tool-schema sizes, repeated conversation-context transmission, step counts, retries, and context-window settings; even a prompt asking only for “OK” used 10,898 tokens in dsh and 13,892 in Hermes. It provides configuration and measurement guidance for reproducing the test with a shared OpenAI-compatible endpoint, emphasizing explicit context settings, compatible tool support, machine-readable usage logs, and provider billing as the ultimate source of truth. The comparison also notes that dsh is positioned as a leaner coding-oriented developer preview with session replay but no persistent memory or messaging integrations, whereas Hermes offers long-term memory, reusable skills, scheduling, chat channels, and dashboards at substantially higher measured token use, suggesting that some users may combine Hermes for persistent coordination with dsh for coding tasks.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.