Skip to content
← articles
updated Hermes AgentAI AgentsToken CostHarness EngineeringSelf-Hosted AI

How Much Does Hermes Agent Cost? Where the Tokens Actually Go

Hermes Agent is free and MIT-licensed. The real cost is tokens and the machine it runs on, and both move almost entirely with how you configure the runtime: skills, memory, compression, tool search, the background review loop.

Hermes Agent costs nothing to install. Nous Research ships it MIT-licensed, and the install script never asks for a card. That answers one question and tells you nothing about the one you actually care about.

I run Hermes in production: a personal assistant on Telegram, and the thirteen agents in my Paperclip org, all on a VPS I pay for by the month. The software never sent me an invoice. The model provider is where the money goes, and every big swing I measured came from configuration, not from the model. This is where that money actually goes, and which keys in ~/.hermes/config.yaml move it.

Is Hermes Agent free?

Yes, the software is. Hermes Agent is open source and MIT-licensed: no seat price, no feature gate, no tier that unlocks the learning loop for a fee. In that sense it is free the same way Linux or Postgres is free.

What is not free is everything Hermes does on your behalf once it is running: tokens billed by whatever model provider you point it at (or the electricity, if you run a local model), and the machine the process lives on, whether that is a laptop you already own, a VPS, or a serverless sandbox billed by the second. The license is fixed. The bill is not, and the bill is set by configuration, not by Nous Research.

Hermes Agent pricing: free software, a configured bill

Hermes Agent has no pricing page because it has no product to price. What you actually pay for breaks into two buckets: model tokens, metered by your provider or by a subscription you already hold, and compute, for whatever keeps the process alive, cron jobs included, if you run a gateway around the clock.

Both numbers are set by a handful of settings in config.yaml, not by the project. A Hermes instance with a lean context, a trimmed skill set and review routed to a cheap auxiliary model can cost a fraction of the same instance with every default left untouched, on the exact same task, on the exact same model.

Where the tokens actually go: the real drivers of Hermes Agent cost

The honest answer to “how much does Hermes Agent cost” is a map of what loads into the system prompt on every single turn, plus what the learning loop does between turns. Most of it never shows up unless you go read the source.

What is in the prompt, and what controls it

This is the part of the bill you pay before the agent has done any work.
What loadsWhenThe lever
SOUL.md (identity)Every turn, from HERMES_HOMEKeep it short. It is a fixed cost no matter how small the task is.
Project context fileIn the system prompt for the whole session: .hermes.md/HERMES.md, else an AGENTS.md chain walked to the git root, else CLAUDE.md, else .cursorrulesRun each agent from its own scratch directory, not someone else's repo.
Skill indexEvery turn: name and description onlyGrows by one line for every skill installed. Prune skills you never use.
Skill bodiesOnly when the skills tool is actually calledNever let a full SKILL.md leak into the always-on index.
Memory (MEMORY.md, USER.md)Every turn, from HERMES_HOMECurate it like a config file. The model rereads all of it on every wake.
This is the part of the bill you pay before the agent has done any work.

That table is the static cost. The dynamic cost comes from four mechanisms that fire while the agent is working, and each one has a specific, non-obvious trigger.

Background review is the one to watch, and it is also the headline feature. Hermes nudges a memory review every 10 turns and a skill review every 10 tool iterations, both configurable. On the same model, the review runs as a forked agent that replays the entire session history and leans on the warm prompt cache to keep that affordable. Point review at a different, cheaper model through auxiliary.background_review and the behavior changes: instead of a full replay, the fork gets a digest, the most recent 24 messages kept verbatim plus everything older collapsed into one summary turn. Same-model review is simple and cache-cheap. Cross-model review is cheaper per call and loses resolution on anything older than 24 messages. Pick based on what you are optimizing for, not on the default.

Compression is the safety valve, not a cost itself. The main compression pass triggers once the conversation hits 50 percent of the model’s context window by default, collapsing older turns so the agent does not stall out against the limit. A separate micro-compaction pass exists in the code but ships off. If your sessions run long enough to compress often, that is a sign the earlier layers, skills and memory, are overloaded, not a reason to fight the compressor.

Tool search is the one most people misunderstand. It does not turn on at some percentage of your context window. It turns on the moment any MCP server or plugin tool is present at all, full stop. What the 5 percent default actually governs is the size of the listing it embeds in the prompt once it is on: a full listing, then names only, then nothing, as the available budget tightens. If you have a handful of MCP tools, tool search barely costs you anything. If you have over a hundred, as my fleet did at one point, it is the difference between an unreadable wall of tool schemas and a short list the model can search.

Programmatic tool calling is the one that saves you the most when a task fans out. Instead of calling a tool fifty times and paying for fifty results in the context, the model writes a script that calls Hermes tools over RPC, and only the script’s stdout comes back, capped at 50 KB with the head and tail kept. Everything past that cap spills to a file on disk, capped at 5 MB, entirely out of the context window. A loop over a hundred files costs one summary line instead of a hundred tool results.

What I measured running it myself

Numbers from my own runs, not the docs

−22%
input tokens per wake9,044 → 6,992, one scratch dir per agent
6×
cost of a 'lighter' template4.6k vs 28.6k tokens, and it broke a code path
14-15k
tokens of skill bodiesinjected on every wake, before the fix
All three measured running Hermes inside my Paperclip org, in June. Hermes has shipped many releases since; treat the mechanism as current, the exact figures as a snapshot.

The first number came from a leak I did not expect. My Paperclip agents ran inside Paperclip’s own application folder, so Hermes auto-discovered Paperclip’s internal AGENTS.md, written for a different program entirely, and loaded it on every wake. Giving each agent its own scratch working directory took input tokens from 9,044 to 6,992 per wake, a 22 percent cut, documented in my adapter’s README.

The second is a trap I walked straight into and wrote up in the Lab: a “LIGHT” prompt template that I assumed would be cheaper cost six times more than the full one, 4.6k input tokens against 28.6k, and it broke a code path on top of costing more. I had never measured it. I had assumed. The write-up is Measure, don’t assume, and it is the single habit that would have saved me the most money across this whole project.

The third came from the same adapter path: injecting full skill bodies into the prompt on every wake, instead of the frontmatter-only index Hermes’ own skills tool is built for, cost 14,000 to 15,000 tokens a call. Switching to index-only and loading bodies on demand through the skills tool fixed it.

Subscriptions vs API keys

In June, part of my fleet’s token budget came from gpt-5.4 through Codex, riding on a ChatGPT Plus subscription. Subscriptions are tempting for an agent fleet, and provider terms for third-party apps change on their own schedule. Check the current terms before you plan a budget around one.

My back-of-envelope estimate at the time, for running the fleet’s thinking agents on GLM-5.1 through OpenRouter: about $0.05 per run, somewhere between $5 and $15 a month for the whole operation. Compare that to a $100-a-month Max plan bought for the same volume of work, and the gap is the entire argument for metering tokens directly instead of defaulting to the subscription that is already in your wallet. Treat that number as my estimate from one point in time, not an audited figure, and run your own math before you copy it.

How much does Hermes Agent cost compared to OpenClaw?

Both runtimes put the real cost in a loop that fires on its own, and both loops are a setting, not a fixed property of the project. OpenClaw wakes on a heartbeat by default, every 30 minutes, and each beat is a full agent turn. Its own docs put an untuned heartbeat at roughly 100,000 tokens per run, dropping to 2,000 to 5,000 tokens with isolatedSession turned on, because that setting stops resending the full conversation history on every beat.

Hermes has no ambient heartbeat. It only wakes on cron or when something talks to it, so it does not carry that specific cost. What it carries instead is the background review described above, which fires on turns and tool iterations rather than a clock. Neither comparison is “Hermes is cheaper” or “OpenClaw is cheaper.” Both bills come from a loop you can tune. OpenClaw documents its number: one setting takes a beat from about 100k tokens to 2k to 5k. Hermes publishes no number for its review loop, so route it to a cheaper model and measure your own sessions.

The cheapest settings that matter most

Before you add a bigger model to the budget

  • Required:
    Give every agent its own scratch working directory.Stops auto-discovery from loading an AGENTS.md written for a different program. Cost me 22 percent of every wake until I fixed it.
  • Required:
    Keep skill bodies out of the always-on prompt.The index costs little. A full SKILL.md injected on every wake cost 14 to 15k tokens in my setup.
  • Required:
    Route background review to a cheaper auxiliary model.auxiliary.background_review swaps the same-model full-history replay for a 24-message digest plus a summary.
  • Required:
    Measure before you tune compression or tool search thresholds.The 50 percent compression trigger and the 5 percent tool search budget are reasonable defaults, not universal ones. Change them after you have watched a real session, not before.
  • Optional:
    Let programmatic tool calling carry the fan-out work.A script that loops over many items returns one capped stdout summary instead of one tool result per item.
  • Optional:
    Set the model in config.yaml, never in whatever wakes Hermes.The runtime owns the model and the reasoning effort. Every time an orchestrator tried to own it instead, the persona or the bill drifted.
None of these require a different model. All six are configuration you already have access to.

The token bill on any agent runtime is a plumbing problem before it is a model problem. The full method for finding leaks like these, independent of which runtime you run, is in cutting your AI coding agent’s token cost.

Hermes Agent cost, quick answers

Is Hermes Agent free?

The software is: open source and MIT-licensed, with no paid tier. What costs money is the model tokens Hermes uses and the machine it runs on, and both are set by how you configure the runtime, not by the project.

How much does Hermes Agent cost?

There is no fixed number. It depends on the model you point it at, how many tokens your sessions use, and how aggressively the background review, skill bodies and context files load on every wake.

My estimate in June, for running my fleet's thinking agents on GLM-5.1 through OpenRouter, was about $0.05 per run and $5 to 15 a month. Treat that as one estimate from one setup, not a quote.

What drives Hermes Agent pricing the most?

The background review loop and whatever auto-loads into the system prompt on every turn: SOUL.md, the one project context file, the skill index and memory. A single misconfigured context file cost 22 percent of every wake in my own setup.

Does Hermes Agent cost more than OpenClaw?

Neither is inherently cheaper. Both put their biggest variable cost in a loop that runs on its own: OpenClaw's heartbeat, Hermes' background review. OpenClaw documents a drop from about 100k to 2k to 5k tokens per heartbeat with isolatedSession. For Hermes, route the review to a cheaper model and measure.

Can I run Hermes Agent for free with a local model?

Yes. Point config.yaml's provider at custom and a local endpoint such as Ollama, and the only cost left is the machine's electricity and whatever that model's accuracy costs you in retries.