Skip to content
← articles
updated Hermes AgentMCPAI AgentsTool UseSelf-Hosted AI

Hermes Agent and MCP: Wiring Servers That Actually Show Up

How MCP servers are configured in Hermes Agent, what changed with the v0.21.5 Connectors page, how tool search really defers tools, and the production fixes a cold-spawn race and a hard-scope flag forced on me.

An MCP server connecting is not the same as an agent using it. I learned that the expensive way: 157 tools wired, zero tool calls, and a model that swore the tools did not exist. The tools existed. The config was fine. What was missing was timing, and nothing in a quickstart warns you about timing.

This is the config, the current UI, and the production fixes for running MCP servers against Hermes Agent, checked against v0.21.5, commit 9c6fe3aa.

MCP servers in Hermes Agent start in one block

Every MCP server Hermes talks to lives under mcp_servers in ~/.hermes/config.yaml, one entry per server, and the entry’s shape depends on whether the server speaks stdio or HTTP.

mcp_servers:
  project_fs:
    command: "npx"
    args:
      ["-y", "@modelcontextprotocol/server-filesystem", "/home/user/project"]
    tools:
      include: [read_file, list_directory]

  github:
    url: "https://api.githubcopilot.com/mcp/"
    headers:
      Authorization: "Bearer ${env:GITHUB_TOKEN}"
    tools:
      exclude: [delete_repo]

A stdio server gets command, args and env. An HTTP server gets url and headers instead, plus TLS keys (ssl_verify, client_cert, client_key) when the endpoint needs mTLS. Both shapes share enabled, timeout, connect_timeout, trust and a tools policy with include/exclude globs plus resources/prompts toggles for the utility wrappers Hermes adds on top. ${VAR} and the Cursor-style ${env:VAR} both resolve from the active profile’s .env, so a snippet copied from a Cursor or Claude Code config works unchanged.

Two keys matter more than the quickstart lets on. lazy: true registers a server’s tools from the on-disk schema cache at startup and only spawns or connects it on the first real call, which is the right default for a server you rarely use. idle_timeout_seconds and max_lifetime_seconds recycle a stdio server after it sits idle or gets old, which matters for any server that leaks memory slowly.

You do not have to hand-edit YAML for the common case. hermes mcp add writes the entry for you, and hermes mcp test <server> connects and lists what it discovers:

Add and verify a server

  1. Local stdio server

    hermes mcp add github --command npx --args -y @modelcontextprotocol/server-github
  2. Verify it connects

    hermes mcp test github

Inside a running session, /reload-mcp picks up a config change without a restart. That slash command is the fix for a bug that cost me a restart in June: an MCP server added after the gateway had already booted would connect, but its toolset never registered for activation until the gateway restarted. /reload-mcp is the documented remedy today. If you are running Hermes from before that slash command existed, budget for a restart every time you touch mcp_servers.

The MCP command center replaced the old MCP tab

Hermes Desktop used to have an MCP tab. As of v0.21.5 it has a Connectors page instead, and the change is more than cosmetic. The release rolled in a catalog of curated, Nous-approved servers, health checks per connection, and “Connect now” prompts that fire the moment you install a plugin that ships its own MCP servers, so a freshly installed plugin’s tools and skills go live in every chat you already have open, not just new ones. The same release window added mcp.discovery_concurrency to cap how many servers Hermes probes at once during startup discovery, which is the kind of knob you only need once you have enough servers for discovery itself to become a bottleneck.

Tool search defers every MCP tool, the budget is the only variable

The moment any MCP or plugin tool exists in a session, Hermes moves it behind a search tool instead of listing it in the system prompt. This is not gated by how many tools you have. tool_search.py activates whenever a deferrable tool exists at all, full stop. There is no “under 10 percent of context, list everything” threshold in the code.

What the config actually controls is the size of the compact listing Hermes still embeds for the deferred tools: threshold_pct, 5.0 by default, capped at listing_max_tokens, 4,000 tokens by default. The budget is min(listing_max_tokens, threshold_pct% of the context window), and an unknown context length falls back to a flat 10,000-token leg, which is 5 percent of a typical 200K window. Raise threshold_pct and you get a richer embedded listing at the cost of prompt space on every turn. Lower it and the model relies more on the search tool itself.

This is the mechanism that makes hard-scoping an agent’s tools a much smaller win than it looks. My fleet had 114 MCP tools sitting behind search in June. The token cost of “too many tools” was already mitigated by deferral; the only thing hard-scoping bought me was risk. More on that below.

hermes mcp serve: Hermes as a source for other agents

The config above is Hermes as an MCP client, pulling tools in. Hermes can also run as a server. hermes mcp serve starts a stdio MCP server that exposes your messaging conversations as tools: list conversations, read history, send messages, poll live events, manage approvals. Point Claude Code, Cursor or Codex at it:

{
  "mcpServers": {
    "hermes": { "command": "hermes", "args": ["mcp", "serve"] }
  }
}

The module’s own docstring says it matches OpenClaw’s 9-tool channel bridge surface, which tells you where the design pressure came from: once one self-hosted runtime exposes its channels to a coding agent over MCP, the other one has to.

Hermes Agent MCP in production: the fixes that matter

Everything above is the documentation. Everything below is what the documentation does not warn you about, learned running a Paperclip agent fleet on Hermes since June.

What MCP in production actually taught me

  1. June157 tools, zero calls

    Every agent had its full toolset wired through an aggregating gateway in front of 7 backend MCP servers. They called none of it and reported the tools did not exist. The real cause lived in agent.log: the gateway cold-spawned its 7 backends per session, 4.5 to 27 seconds, past Hermes' own sub-second discovery window. A warm backend answers tools/list in about 10 milliseconds. The agents just never waited that long.

  2. JuneControl plane and data plane are different jobs

    The fix was not to abandon the aggregating gateway, it earns its keep as an admin catalog. The fix was to stop routing hot traffic through it. mcp-pooler holds one persistent warm session to the gateway, caches tools/list, and proxies tools/call, so every short-lived agent session sees discovery in about 50 milliseconds instead of tens of seconds. In config.yaml that means one HTTP entry pointed at the pooler, not one entry per backend.

  3. Junehermes mcp test proves less than it sounds like

    The command connects and lists tools, which is a static check: it proves spawn plus tools/list, nothing about auth or reachability under load. Real proof needed a tools/call plus a direct API probe. Today's docs describe the same three exit codes, 0 connected, 1 failed, 3 unconfigured, which still means the same limited thing.

  4. JuneThe hard-scope flag replaces, it does not narrow

    Scoping an agent with Hermes' toolset flag replaced the whole base toolset instead of adding to it, which stripped the agent's terminal and skills and left it a zombie with hands cut off. Combined with tool search already paying for the token cost of a large toolset, hard-scoping stopped being worth the risk. I moved to soft scope: doctrine in a TOOLS.md plus a project fence, and it held. An agent with the full pool simply never touched the namespace it was told to leave alone.

  5. JuneValidate the consumer before you build the source

    I spent an afternoon building four hard-scoped MCP namespaces before checking how the toolset flag actually consumed them. The whole afternoon's work got deleted an hour after it passed its own tests, because I had proven the wrong thing correctly.

  6. OngoingThe agent is an unreliable narrator

    Every MCP failure I debugged by believing the agent's own account cost me time. agent.log was right every time the agent's story was wrong. When a tool call fails, read the log before you read the chat transcript.

The June lessons predate v0.21. Read them as history, not as today's bug list.

The cold-spawn race and the scoping mistake are both written up in full in the Lab: 157 tools and zero calls and validate the consumption before you build the source. The data MCPs behind that fleet were Ubersuggest, GA4 and Search Console, catalogued in the aggregating gateway and reached through the pooler. None of it was exotic. Every failure traced back to a race, a flag’s real semantics, or a sequencing mistake, never to the model.

Words that get you proof, not a narration

  1. 01

    Make the agent call the tool, not just name it.

    Tool search and hermes mcp test both stop at listing. A tool that shows up in a search result has not been proven to return anything.

    Type this

    Call that tool once with a real argument right now and paste the raw result. Don't tell me it's available, show me it returning data.
  2. 02

    When it says a tool does not exist, ask for the log line, not the excuse.

    The agent narrates a failure. agent.log records it. The two disagree more often than you would guess.

    Instead of

    Why didn't that tool work?

    Type this

    Open agent.log and quote the exact line for the last failed MCP call. Don't summarize it, quote it.
  3. 03

    Before you trust a hard scope, ask what it actually removed.

    Hermes' toolset flag replaces the base toolset instead of narrowing it. An agent that lost its terminal often will not volunteer that.

    Type this

    List every tool you have access to right now, then tell me what you had before I added the -t flag.
Three sentences that turn a confident agent report into something you can actually check.

Before you wire a new MCP server into Hermes

  • Required:
    Decide stdio or HTTP first, the keys do not mix.command/args/env for a local process, url/headers for a remote endpoint. One shape per server.
  • Required:
    Reach for tools.include before you reach for a scoping flag.Tool search already pays most of the token cost of a big toolset. A broad agent fenced by doctrine is safer than a hard-scoped one missing its own terminal.
  • Required:
    Treat hermes mcp test as a smoke test, not a health check.It proves the server spawns and lists tools. Prove a real tools/call separately before you trust the wiring.
  • Required:
    If a server sits behind an aggregating gateway, put a warm pooler in front of it.Ephemeral agent sessions do not survive a 4 to 27 second cold spawn. A cached tools/list and one warm upstream session do.
  • Required:
    When an agent says a tool does not exist, open agent.log before you believe it.The agent narrates. The log is ground truth.
Everything here came from a production fleet, not a demo.

Hermes Agent MCP, quick answers

How do I add an MCP server to Hermes Agent?

Add an entry under mcp_servers in ~/.hermes/config.yaml, or run hermes mcp add <name> --command ... --args ... for a stdio server (use --url for HTTP instead of a command). Verify with hermes mcp test <name>, then reload a running session with /reload-mcp instead of restarting.

What is the MCP command center in Hermes Agent?

As of v0.21.5, Hermes Desktop's old MCP tab is the Connectors page: a catalog of curated servers, per-connection health checks, and a 'Connect now' flow that activates a freshly installed plugin's MCP tools in every chat already open.

Does Hermes Agent's tool search only activate above a token threshold?

No. Tool search activates the moment any MCP or plugin tool exists in the session, regardless of how many tools that is. The only configurable thing is the size of the embedded listing for deferred tools, 5 percent of the context window by default, capped at 4,000 tokens.

Can Claude Code use Hermes Agent as an MCP server?

Yes. hermes mcp serve exposes your messaging conversations, list, read history, send, poll events, manage approvals, as MCP tools any client can call, Claude Code, Cursor or Codex included.

Why does my Hermes agent say its MCP tools do not exist?

Usually a discovery race, not a missing tool. If the server sits behind an aggregating gateway that cold-spawns its backends per session, discovery can take seconds while Hermes' own window is under a second. Read agent.log before you trust the agent's explanation, and put a warm pooler in front of the gateway if cold starts are the cause.