Is Hermes Agent Safe? The Agent Should Never Hold the Token
An always-on agent with a shell is a colleague with root. What actually went wrong running Hermes Agent in production, the controls that caught it, how OpenClaw compares, and a checklist for before you give any agent a shell.
An always-on agent with a shell is a colleague with root, one who never sleeps, reads everything you point it at, and will cheerfully print a secret to a log if nothing stops it. Security for that colleague is mostly about two questions: what can it reach, and what can it print. The model behind it matters far less than either.
I run Hermes Agent in production on my own VPS. Every incident below actually happened to me. Every control below I checked in the source, not the marketing page.
Is Hermes Agent safe?
It has real controls, and its own code is explicit about which of them are not a security boundary. Approval prompts gate risky commands on every messaging surface, an emergency-stop file can halt all new work, and an egress proxy keeps credentials out of the Docker sandbox. None of that makes Hermes safe by default. It makes Hermes safe to configure correctly, which is a different claim, and the gap between the two is where every incident below happened.
Treat it like what it is: a colleague with a shell, not a sandboxed toy. Give it its own machine or container, scoped credentials, and never a key you cannot rotate.
What actually went wrong when I ran it
None of this is theoretical. Four separate incidents, four separate lessons, running Hermes and Paperclip in production since June.
Incident, cause, fix
| What happened | Root cause | Fix |
|---|---|---|
| API tokens showed up in a session log | A dump of config.yaml inside a Hermes session printed them straight into the log | Rotated the tokens |
| The web dashboard hung on connecting | Chrome does not resend basic auth on a WebSocket upgrade. The same trap hit OpenClaw's own UI. | A second Traefik router with a token, separate from basic auth |
| A run failed with permission denied on agent.log | My own diagnostics, run as root inside the container, left root-owned files in .hermes | chown back to the service user |
| A provider token landed on a service that never used it | My auth tooling propagates tokens to every managed service, but each app decides for itself which env var it reads | Treat propagation and consumption as two separate questions, not one |
The pattern across all four: the agent did not leak anything by being clever. A tool printed what it was told to print, a browser dropped a header it was never going to resend, a diagnostic ran with the wrong user, and a token reached a box that had no use for it. Agent security mostly fails at the plumbing, not at the model deciding to misbehave.
The second-order lesson from the WebSocket fix is worth sitting with. The bypass route used a token in the URL query string, ?token=, which is itself a smaller version of the same problem: anything in a URL ends up in access logs, browser history and referrer headers. A fix for one leak can quietly open a narrower one. Check what you log before you ship the patch, not after.
Hermes agent security controls, from the source
Four mechanisms do the real work, and Hermes’ own source is explicit about which one is cosmetic.
What stands between the agent and the damage
- 1
Approval on every surface
A risky terminal command stops and asks, with the same wording on Telegram, the CLI or any other gateway: 'Hermes wants to run a command that needs your OK,' with the reason it was flagged and a timeout after which the command does not run.
- 2
ESTOP, a sentinel file
hermes pause writes a file at $HERMES_HOME/ESTOP. While it exists, the cron scheduler, kanban dispatcher and new gateway turns skip all new work. In-flight work is never killed. hermes resume removes it.
- 3
An egress proxy for the Docker backend
Hermes' own code calls it a 'host-side credential firewall.' It builds the mount, env and host arguments that route a sandbox through it, and it guards the exact config surfaces (forwarded env, extra args) that could otherwise weaken or bypass it.
- 4
Output redaction
A dedicated module strips sensitive strings out of what the agent prints, across the approval prompt, the terminal tool and delivery to a gateway. It reduces the blast radius of a leak. It does not prevent one.
Every guard here is defense-in-depth, NOT a security boundary: the terminal tool runs as the same OS user and can read/write anything.
That line is from agent/file_safety.py in the Hermes source, not from a blog post about Hermes. It is the most honest sentence in the whole codebase. A file guard stops a model that respects a tool error. It does nothing against a model that reaches for bash instead, because the terminal tool is the same OS user either way.
Two more structural controls matter here. Hermes profiles give you “multiple isolated Hermes instances,” each with its own HERMES_HOME, config and state database, which is how you keep one compromised agent from reading another agent’s memory or credentials on the same box. And the execution backend you pick (local by default, or Docker, SSH, Singularity, Modal, Daytona or Vercel Sandbox) is the actual blast-radius control: an agent that runs its commands in a disposable sandbox loses far less than one that runs them on the box it lives on.
Credentials behind a proxy, never in the agent’s hands
The agent should never hold a token it could print. The pattern that works is a proxy that holds the credential and a config that only ever points the agent at the proxy.
My own mcp-pooler does exactly this for MCP traffic. It holds one upstream bearer token, set once as an environment variable on the pooler itself, and every short-lived agent downstream talks to the pooler’s own endpoint, which takes no credential at all. The agent cannot leak the upstream token because the agent never receives it. The tradeoff is explicit in the pooler’s own README: that downstream endpoint has no auth of its own and is meant to run on a private network, not exposed to the internet. A proxy that hides one credential is not a reason to skip network-level isolation.
The same logic applies to every source an agent reads, not just the gateway in front of them. Scope Slack and Google Workspace MCP sources to read-only wherever the tool list allows it, and treat the OAuth client or bot token behind each one as something you issued and can revoke, never something borrowed from another system. A Git source is the same question again: a token scoped to one repo leaks a lot less than one scoped to your whole account.
Is OpenClaw safe? Same question, different defaults
OpenClaw and Hermes start from the same default: commands run on the host. Hermes ships with the local backend selected, and OpenClaw runs tools on the host unless you turn sandboxing on. OpenClaw’s own documentation calls that choice out directly, offering Docker, Podman, SSH, OpenShell or Crabbox as opt-in sandboxes, and stating plainly that even the sandbox is “not a perfect security boundary.” Its heartbeat wakes the agent on its own every 30 minutes by default, which is a feature for a personal assistant and a wider attack window for anything you did not mean to give it.
OpenClaw security, like Hermes agent security, is a configuration decision, not a property either project ships for you. Exec approvals exist on both. Neither sandboxes by default; both let you pick a sandbox when you configure it. Neither fact tells you whether a specific deployment is safe. The questions in the checklist below do.
AI agent security checklist: before you give an agent a shell
OWASP’s GenAI Security Project names the pattern directly: LLM06:2025, Excessive Agency, the risk that an agent is granted more permission, more tools or more autonomy than the task in front of it needs. Every incident in this article is a flavor of that same risk, not a flaw unique to one runtime.
Before you give an agent a shell
- Required:It runs as its own user, in its own sandbox or container.File-write guards are not a security boundary. The OS user the tool runs as is the real one.
- Required:Every credential it touches sits behind a proxy or a scoped token you issued yourself.The agent should never hold a secret it could print. If it must, that secret should be the smallest one that works.
- Required:Risky actions stop and ask, on whatever surface you actually read.An approval prompt nobody sees is not an approval gate. Match the surface to where you are.
- Required:You have a kill switch that stops new work without killing work in flight.Hermes' ESTOP file is one shape of this. Know yours before you need it.
- Required:Publishing, sending and deleting go through a human, by design, not by accident.My own agents never got write access to the channels I publish to. They produce. I publish.
- Required:You know what it logs, and you have read a log file, not just the agent's own account of what happened.A config dump, a URL query string, a stack trace: all of them are places a secret hides in plain sight.
Hermes Agent security, quick answers
Is Hermes Agent safe?
It ships real controls: approval prompts on every messaging surface, an ESTOP file that halts new work, and an egress proxy that keeps credentials out of the Docker sandbox. Its own code states that file-write guards are defense-in-depth, not a security boundary.
Whether a specific deployment is safe depends on the execution backend, the credentials it can reach, and whether risky actions require a human. Those are configuration choices, not defaults.
What is Hermes agent security actually protecting against?
Mostly plumbing failures, not a model deciding to misbehave: a tool printing a secret it was told to print, a diagnostic run with the wrong OS user, a token reaching a service that never needed it. Every incident I hit in production was one of these, not the model going rogue.
Is OpenClaw safe?
Like Hermes, OpenClaw runs commands on the host by default and makes sandboxing opt-in, across Docker, Podman, SSH, OpenShell or Crabbox. Its own documentation says the sandbox is not a perfect security boundary. That is an honest default to know about, not a verdict on any specific setup.
How does Hermes Agent keep credentials away from the agent?
Through a proxy pattern: a component like mcp-pooler holds the upstream bearer token and the agent only ever talks to the proxy's own endpoint, which carries no credential. Combine that with read-only scopes on sources like Slack and Google Workspace wherever the tool list allows it.
What is AI agent security, in practice?
Two questions, repeated for every tool and every credential: what can this agent reach, and what can it print. OWASP's GenAI Security Project calls the general failure mode Excessive Agency. Everything else, from sandboxing to approval prompts to redaction, is a specific answer to one of those two questions.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)