Skip to content
← articles
updated Hermes AgentAI AgentsHarness EngineeringNous ResearchSelf-Hosted AI

What Is Hermes Agent? A Developer's Guide From Running It in Production

Hermes Agent is Nous Research's open-source, self-hosted agent runtime with memory, a learning loop and more than 25 messaging gateways. How it works under the hood, what the learning loop really does, where it breaks in production, and when a developer should pick something else.

Search for Hermes Agent and you get install guides. Paste a curl, pick a model, connect Telegram, done. That is the easy part, and it is the part that matters least once the thing has been running for a month.

I run Hermes in production. Thirteen agents in a Paperclip org run on it, a separate Hermes on Telegram is my personal assistant, and at my day job another one writes me a daily briefing from the company’s MCP servers. I published an adapter for it, sent a fix upstream, and read its source more times than I wanted to. This is the guide I wanted when I started: what Hermes is under the hood, what the learning loop actually does, where it breaks, and when a developer should pick something else.

What is Hermes Agent?

Hermes Agent is an open-source, MIT-licensed agent runtime from Nous Research that you host yourself. You reach it from a terminal, a desktop app, a web dashboard, an OpenAI-compatible API, or more than 25 messaging platforms. It acts through a shell, files, a browser and code execution. It keeps memory across sessions, searches its own history, and writes its own skills after hard tasks.

Three things separate it from the coding agent you already use. It runs without you: on a server, on a schedule, answering a message on your phone. It remembers: a memory file, a profile of you and a searchable history survive every session. And it learns: every few turns, a background review decides whether something it just did is worth saving.

Hermes Agent at v0.21.5

250k
GitHub starsand nearly 48k open issues and PRs
38
model provider pluginsplus any OpenAI-compatible endpoint
7
execution backendslocal to serverless sandboxes
25+
messaging platformsTelegram to Matrix to WeChat
The breadth is real. So is the issue count. Both come from a project moving very fast.

How Hermes Agent works

Strip Hermes down and you get the same four parts as any harness: a loop, tools, context management and controls. What makes it Hermes is how much it put in each layer.

Hermes Agent, from the outside in

  1. 1

    Gateways

    CLI, a terminal UI, a desktop app, a web dashboard, an OpenAI-compatible API server, ACP for editors like Zed, and more than 25 messaging platforms. Every surface talks to the same agent.

  2. 2

    Context

    A system prompt built in tiers: a stable identity from SOUL.md, one project context file, a compact index of skills, and memory. Compression kicks in at half the context window by default.

  3. 3

    Tools

    Terminal, files, browser, web, vision, MCP servers, plus three primitives that change what a run costs: programmatic tool calling, tool search and subagent delegation.

  4. 4

    Execution backends

    Where commands actually run: local, Docker, SSH, Singularity, Modal, Daytona or Vercel Sandbox.

  5. 5

    Learning

    A background review that saves memories and skills, a weekly curator that archives what goes unused, and full-text search over every past session.

Model providers plug in at the side: 38 provider plugins, and local models through any OpenAI-compatible endpoint.

Nous Research builds models, and it shows. Hermes ships a batch runner, a trajectory compressor and a SWE runner that write tool-calling trajectories in Hermes’ own format. The runtime doubles as a way to generate training data for the lab’s models. That is a structural reason for the project to keep existing, and it explains why the loop primitives are deeper than the consumer features.

What Hermes reads before it answers

Most surprises I had with Hermes came from not knowing what it put in front of the model. Here is the map. How I feed it a second brain without letting it write there is in Hermes Agent and a second brain.

What feeds the prompt

  • ~/.hermes/// HERMES_HOME, one per profile
    • config.yamlyou// model, provider, reasoning effort, MCP servers
    • SOUL.mdalways// identity, loaded on every turn
    • MEMORY.mdlearned// what it saved about the work
    • USER.mdlearned// what it saved about you
    • skills/// SKILL.md folders, agentskills.io format
    • state.db// every session, searchable with SQLite FTS5
  • your working directory/// only one context file loads, first found wins
    • .hermes.md// 1st
    • AGENTS.md// 2nd, walked from the git root
    • CLAUDE.md// 3rd
    • .cursorrules// 4th

Two lessons cost me real money here.

Auto-discovery is a feature until the working directory belongs to someone else. My Paperclip agents ran inside Paperclip’s own app folder, so Hermes found Paperclip’s internal AGENTS.md and loaded it on every wake. That was 22 percent of each wake’s input tokens, for instructions written for a different program. A scratch directory per agent took input from 9,044 tokens to 6,992.

Identity goes in SOUL.md, nowhere else. Hermes opens its system prompt by introducing itself as Hermes Agent. When my persona arrived later as a user message, the model got two identities and picked the safe average. On my own offer-design test, GLM-5.1 scored 9.55 through OpenCode and 7.0 to 7.9 through Hermes with the persona in the wrong place. Same model. The prompt layout was the variable.

The learning loop, without the marketing

The loop is the headline feature and the one people describe worst. Here is what the code does.

Every 10 turns Hermes nudges a memory review, and every 10 tool iterations it nudges a skill review. Both intervals are configurable. The review runs as a forked agent that can only touch the memory and skill tools, and it decides whether something from the session is worth keeping. A curator runs weekly when the agent is idle, marks unused skills stale after 14 days and archives them after 30. It never deletes, so anything it archived can come back.

My take after months of running it: the loop is good at procedure and bad at judgment. It will remember the flag a tool needs, the order a deploy goes in, the workaround for a flaky API. It should not be the thing that rewrites how an agent thinks. My persona skills carry judgment, and I write those by hand from material I studied. A background fork that saw one session is not the right author for them. The full split is in Hermes Agent skills.

The tools that change the economics

Three primitives matter more than the long list of built-in tools.

Programmatic tool calling. The model writes a Python script that calls Hermes tools over RPC, and only the script’s stdout comes back into the context, capped at 50 KB with the head and tail kept. Everything else spills to a file on disk. If your agent loops over fifty items calling a tool on each, this is the difference between fifty tool results in the context and one summary line.

Tool search. When any MCP or plugin tool is present, Hermes moves those tools behind a search tool and embeds a compact listing sized to a small budget, 5 percent of the window by default. My fleet had 114 MCP tools behind search. With deferral, the cost of “too many tools” mostly disappears, and hard-scoping each agent’s tools stops being worth the risk of breaking it.

Subagents. delegate_task forks isolated children, up to 10 at once by default, and children can delegate again within a depth budget. Since v0.21 you can list, steer and stop subagents while they run, and validate what they return against a JSON schema.

There is more: a Mixture-of-Agents mode that aggregates answers from several models, cron jobs you can chain and reply to, and a Bot Mode where named agents share group chats. Useful, but the three above are the ones that change what a run costs. The whole bill is broken down in how much Hermes Agent costs.

Seven places your commands can run

Hermes separates the agent from where its commands execute. The backend list at v0.21.5: local, Docker, SSH, Singularity, Modal, Daytona and Vercel Sandbox. Modal and Daytona give you sandboxes that persist between tasks without a machine you keep paying for.

Run the agent on one box and execute somewhere disposable. The Docker backend ships an egress proxy that keeps credentials out of the sandbox. Risky terminal commands stop and ask for approval on whatever surface you are using, Telegram included, and an ESTOP file in HERMES_HOME pauses all new cron, kanban and gateway work. File-write guards exist too, but the code is explicit that they are not a security boundary: the tool runs as your OS user. The rest is in Is Hermes Agent safe?.

How I run Hermes in production

My setup is boring on purpose. A VPS turned into a Docker Swarm with Traefik and Portainer by bento, my installer. Hermes runs headless in a container with the Telegram gateway on, and that is the assistant I talk to by voice. Paperclip runs on the same swarm, and every agent in the org calls the Hermes binary through my adapter. MCP servers sit in a catalog gateway with mcp-pooler in front of it, so short-lived agents see their tools in about 50 milliseconds instead of several seconds. The full MCP setup is in Hermes Agent and MCP.

The model lives in ~/.hermes/config.yaml, never in the orchestrator. Over the months the fleet ran GLM-5.1 through Z.AI and later DeepSeek V4 Flash through OpenRouter. The narrow agents do fine on cheap models, and how I pick them is in the best model for Hermes Agent. The rule I learned is that the runtime owns the model and the reasoning effort. Every time the orchestrator tried to set them, something drifted.

Where Hermes breaks in production

None of these show up in a demo. All of them showed up within the first month.

Symptom, real cause, fix

What you seeWhat is actually happeningWhat fixed it
The agent says its tools do not existMCP backends cold-spawn past a sub-second discovery windowA warm pooler in front of the MCP gateway
Under an orchestrator, a run that finished its work is marked deadThe orchestrator reads liveness from stdout, and a long silent synthesis looks deadA keepalive on a timer, sent upstream as a pull request
The persona sounds genericThe persona arrives as a user message under Hermes' own identityIdentity in SOUL.md, one source
Every wake costs a fifth more than it shouldAuto-discovery loads another program's AGENTS.md from the working directoryA scratch directory per agent
API keys show up in a session logA dump of config.yaml inside a session wrote them into the logRotate the keys; keep secrets out of anything the agent prints
The dashboard hangs on connectingBrowsers do not resend basic auth on the WebSocket upgradeA token route for the WebSocket, separate from basic auth

The first two are written up in the Lab: 157 tools and zero calls and four minutes of silence that killed an agent. Both taught the same rule. The agent is an unreliable narrator. When it reports a failure, do not debug its story. Read ~/.hermes/logs/agent.log. The log was right every single time.

Then there is the pace. Hermes has nearly 48,000 open issues and pull requests. v0.21.0 shipped regressions in the session database that took six follow-up pull requests and 44 closed issues to settle in v0.21.2. A few open issues I would read before upgrading a production box: a gateway executor with no timeout that freezes the gateway for 120 seconds (#101033), a loop-liveness heartbeat that can be dead from startup (#94758), and a dashboard process that leaks memory until clients get OOM-killed (#80527).

Hermes Agent vs OpenClaw vs Claude Code

They get compared because they all say “agent.” They are built for different jobs.

OpenClaw is a personal assistant that wants to be everywhere you are: native apps for macOS, iOS, Android and Linux, a wake word, on-device voice, a heartbeat that wakes it on its own. Hermes is a runtime that wants to live on a server and run its commands somewhere disposable. In June the difference was depth: Hermes had the learning loop and programmatic tool calling, and OpenClaw did not. By October OpenClaw had shipped both. I picked Hermes for the fleet after reading both codebases, read them again before writing this, and the full comparison is in Hermes Agent vs OpenClaw.

Claude Code is a coding agent you drive inside a repo. I use it every day and it is not a substitute for Hermes, or the other way around. The same person can run both: Claude Code for building, Hermes for the work that should keep happening when the laptop is closed. The longer version is Hermes Agent vs Claude Code.

When not to use Hermes

Pick something else if

  • Required:
    You want a coding agent inside your repo.Claude Code, OpenCode or pi will serve you better. Hermes can code, but its design center is the work that happens when you are not watching.
  • Required:
    You want an assistant on your phone and laptop, with voice.That is OpenClaw's design center: native apps, voice, presence everywhere.
  • Required:
    You will not read Python logs.When Hermes fails, the truth is in agent.log and in the source. The agent's own account is often wrong.
  • Required:
    You need an API that holds still.Releases land about twice a week. Configuration keys, defaults and features move.
  • Required:
    You have one task and no schedule.Memory, cron and gateways pay off when an agent runs for weeks. For one task, a plain session is simpler.
If none of these apply, Hermes is the most capable self-hosted agent runtime I have run.

Hermes Agent, quick answers

What is Hermes Agent?

Hermes Agent is an open-source, MIT-licensed, self-hosted agent runtime from Nous Research. It acts through a shell, files, a browser and code execution, keeps memory across sessions, searches its own history, writes its own skills, and is reachable from a terminal, a desktop app, a web dashboard, an API or more than 25 messaging platforms.

Is Hermes Agent free?

The software is free and MIT-licensed. You pay for the model tokens and for wherever it runs.

The real cost is in how you configure it: what loads into the prompt on every turn, how often the learning loop fires, and which model each job gets.

Is Hermes Agent safe?

It has real controls: approval prompts for risky commands on every surface, an emergency stop file, and an egress proxy that keeps credentials out of the Docker sandbox. Its own code says the file-write guards are not a security boundary.

Treat it like a colleague with a shell. Give it its own machine or sandbox, scoped credentials, and never a key you cannot rotate.

What is the best model for Hermes Agent?

The cheapest one that passes your own test, run through Hermes, not in a plain chat. The runtime changes the result: the same model scored 9.55 through OpenCode and as low as 7.0 through Hermes when my persona landed in the wrong place.

Set the model in ~/.hermes/config.yaml and keep orchestrators out of it.

Does Hermes Agent work with Claude Code?

They do different jobs and coexist well. Hermes speaks ACP for editors such as Zed, and hermes mcp serve exposes its conversations as MCP tools that Claude Code, Cursor or Codex can call. In my setup, Claude Code is where I build and Hermes is what keeps working when I stop.