Hermes Agent Skills: Let It Learn Procedure, Write the Judgment Yourself
Hermes Agent skills explained from production: the SKILL.md format, the skills hub behind hermes skills install, plugins versus skills, and why the background review that writes its own skills should never touch judgment.
Hermes Agent skills are markdown, not magic. A SKILL.md file, a name, a description, and instructions the agent follows when the description matches what you asked for. What makes the topic worth a full guide is not the format. It is that Hermes writes some of these skills itself, on a timer, without you in the loop, and most people never ask what that loop should and should not be trusted to decide.
I run Hermes Agent in production, and I write my own skills by hand for the parts of my work that need a point of view, not just a procedure. This is the guide to both halves: what a Hermes skill actually is, where hermes skills install pulls from, how a plugin differs from a skill, and the line I draw between what the background review is allowed to learn and what I keep writing myself.
What are Hermes Agent skills?
A Hermes Agent skill is an on-demand knowledge document the agent loads when the work calls for it, built on SKILL.md plus optional references/, templates/, scripts/, and assets/ folders. The format is compatible with the open agentskills.io standard, which is the same spec behind Claude skills, so a skill you wrote for one runtime reads correctly in the other.
One skill, on disk
- ~/.hermes/skills/// single source of truth
- mlops/axolotl/// category, then skill
- SKILL.mdrequired// name, description, instructions
- references/// distilled knowledge, loaded on demand
- scripts/// helpers the skill can call
- templates/// output formats
- assets/// supplementary files
- .hub/// skills hub state
- lock.json// source, hash, scan verdict
- audit.log
What actually enters context is the part worth understanding. The system prompt carries a compact index of every skill’s name and description (agent/prompt_builder.py:1338), which the shipped docs put at roughly 3,000 tokens for a full catalog. Nothing past that loads until the agent asks for it through the skills tool (tools/skills_tool.py:701): the full SKILL.md body on a match, and a specific file under references/ only when the skill’s own instructions tell the agent to fetch it by path. Three tiers, each one gated behind the one before it. This is the same progressive disclosure I described for Claude skills, and the mechanics line up almost exactly, because both follow the same open standard.
Hermes Agent skills, from the source
- 3
- loading tiersindex, body, references, each gated behind the last
- 10
- tool iterationsdefault interval before a skill nudge fires
- 14 / 30
- days to stale / archivethe weekly curator never deletes
- 6
- install sourcesofficial, GitHub, skills.sh, ClawHub, LobeHub, browse.sh
Hermes Agent skills hub: what hermes skills install actually pulls from
hermes skills install is one command over several registries, and the hub treats them very differently depending on who vouches for the source. Official optional skills ship inside the Hermes repo itself and install with built-in trust, no scan warning. GitHub installs cover both direct repo paths and a set of default taps: anthropics/skills, openai/skills, huggingface/skills, and NVIDIA/skills, where NVIDIA’s entries additionally carry a signed skill.oms.sig and a governance skill-card.md, a trust signal most community sources do not offer. Past that sit three community registries: Vercel’s skills.sh, the third-party ClawHub marketplace, and browse.sh, Browserbase’s catalog of site-specific browser-automation skills for sites like Airbnb and Amazon. A direct URL install works too, fetching SKILL.md plus whatever it explicitly references.
Find, read, install
Search a registry
hermes skills search kubernetesRead it before you trust it
hermes skills inspect openai/skills/k8sInstall with a security scan
hermes skills install openai/skills/k8s
Every hub install runs through a security scanner for exfiltration, prompt injection, destructive commands, and supply-chain signals before it lands on disk, and the result is recorded with a content hash in skills/.hub/lock.json. Trust levels gate what --force can override: it can push past a caution-level finding on a community source, and it cannot touch a verdict the scanner marked dangerous. That hierarchy is the honest answer to “what are the best Hermes Agent skills to install”: not the ones with the most downloads, the ones whose source sits highest on that trust ladder, and whose scan you actually read before you shipped it.
If you pair the same skills for a recurring job, a bundle (~/.hermes/skill-bundles/<slug>.yaml) groups several skill names under one slash command, so /backend-dev loads review, tests, and the PR workflow in one call instead of three. It is a YAML alias, not an installer: the skills it lists still have to exist first.
Hermes Agent plugins vs skills: tools and hooks, not procedure
A skill teaches the model what to do. A plugin changes what the runtime can do. That is the whole boundary, and it matters because the two get confused constantly in anything self-hosted.
Plugins register through plugins/plugin_loader.py:144, and the surface is narrow on purpose: register_tool adds a new tool the model can call, register_hook wires into the before and after hooks around a provider call (agent/api_request_hooks.py), register_cli_command adds a subcommand, and register_memory_provider is how an external memory backend like Honcho or mem0 plugs in. Model providers, messaging platforms, and execution backends are built the same way, as plugins, not skills.
Skill
- 01Markdown: SKILL.md plus references
- 02Loads on demand, by description match
- 03Teaches a procedure or a method
- 04Portable across agentskills.io runtimes
- 05You, the curator, or the agent can write one
Plugin
- 01Python, registered through plugin_loader.py
- 02Always loaded if enabled, no trigger
- 03Extends tools, hooks, commands, memory providers
- 04Specific to the Hermes runtime
- 05Only a developer ships one
This is also why a skill is the wrong place to fix a missing capability. If the agent needs a tool that does not exist, no amount of SKILL.md instructs it into existing. That is a plugin’s job, and it is a different kind of work than distilling a procedure into markdown.
The background review that writes its own skills
Every 10 tool iterations, Hermes nudges a skill review, a configurable interval that lives at skills.creation_nudge_interval in agent/agent_init.py (default 10, next to the memory nudge on the same default at memory.nudge_interval). The review runs as a forked agent scoped to the skill_manage tool, and it decides whether whatever just happened is worth saving as a reusable procedure: a multi-step workflow it worked out, a dead end it found the way around, a correction you gave it.
A weekly curator (agent/curator.py) runs on an idle-gated schedule, every seven days with at least two idle hours, and marks a skill stale after 14 days unused, archived after 30. It never deletes. Anything it archived can come back if the agent reaches for it again.
skills:
write_approval: true # false = write freely (default) | true = stage every write
Review a staged skill write
List what is waiting
/skills pendingRead the full diff
/skills diff <id>Apply it
/skills approve <id>
The skill_manage tool also runs an advisory linter on every write, and three of its named rules say more about what a skill should be than most style guides do: incident-log-shape flags a body dense with PR or issue numbers, references-sprawl flags more than 60 reference files, and oversized-body flags a SKILL.md past roughly 24,000 characters, because skill_view loads the whole file and it sits in context for the rest of the session. The shipped guidance for what a skill should capture is lessons, not logs: a generalizable rule plus the mechanism behind it, not a narrated incident. I did not write that rule. I agree with it completely, and it is the same instinct behind every persona skill I build by hand.
The best Hermes Agent skills are the ones you can’t install
My persona skills do not come from a registry. They come from studying someone’s actual material: books, talks, transcripts, and then writing down the principles, the frameworks, and the decision rules in my own words, split by the decision each reference answers. I covered the full method for persona skills on Claude Code, and the architecture carries over unchanged, because a markdown skill is runtime-agnostic: the same file works as a Hermes skill, a Paperclip company skill, or a Claude Code skill, because the loading model is the same standard underneath.
One concrete attempt: a ghostwriter skill that extracted a creator’s voice from 152 of their original posts into explicit rules, lowercase, a provocation hook, a closing open loop. Deploying it into the Hermes bot was the plan, and I want to be precise here: it was built and calibrated, and the deploy itself is still ahead of me, not a finished thing I can point to as running.
The content pipeline that is running: in July, two skills, youtube-packaging and social-post-packaging, alongside a Gary Vaynerchuk persona skill, derived roughly 13 pieces of content from one pillar video. It worked well enough to use, and it had a real gap. Packaging fixed the title before the derivation ran, which meant the number that was supposed to be the hook got locked in too early and never made it into the repositioned pieces downstream. That is a sequencing mistake in a creative pipeline, and no background review caught it, because nothing about it looked like a tool error. It looked like the pipeline working.
One infrastructure detail worth knowing if you run skills across both Hermes and Paperclip: Paperclip’s runtime build strips a skill down to just SKILL.md, dropping the references/ folder entirely. My Paperclip adapter symlinks the source directory back in so the references a skill depends on actually survive the trip. If a skill looks fine in Hermes and comes back generic once an orchestrator is in front of it, check whether your references made the crossing.
Hermes itself ships a related idea in /learn: point it at a book, a doc site, or a workflow you just walked it through, and it authors a SKILL.md plus one reference file per chapter or topic, following the same split-by-decision shape I use by hand. It is a good tool for knowledge-base skills, the kind where the source is the authority and the job is organizing it well. It is not a substitute for the part that actually takes judgment: deciding which voice, which frameworks, which anti-patterns belong in a persona skill in the first place. That decision is mine. /learn can format it. It cannot make it.
What the learning loop is good at vs what it should not touch
This is the one rule worth taking from everything above. The loop is good at procedure. It should not be trusted with judgment.
Good at: procedure
- 01The flag a tool needs and the order a deploy goes in
- 02A workaround for a flaky API, found once and worth keeping
- 03Formatting and naming conventions you corrected once
- 04A multi-step workflow it worked out after a dead end
Should not touch: judgment
- 01Whether a persona's voice is right for a piece of content
- 02Whether a packaging decision made upstream was the correct one
- 03What belongs in a SOUL.md, your identity, always on
- 04Anything that needs the material you actually studied, not one session's output
The mechanism explains why. A background review forks after a session and judges that session against itself: the same model, usually, that just did the work. It can tell you a command needed a flag it was missing. It cannot tell you that a content pipeline’s title decision happened in the wrong order, because from inside that one run, the title looked correct. The gap I hit in July was not a bug the loop could have caught. It was a design mistake in the procedure the loop was faithfully saving.
The split I actually run on: skill is method, loaded lazily, one skill per topic. SOUL.md is judgment, always on, one per agent. Same judgment across a set of tasks, one agent plus a shelf of skills it loads as needed. Different judgments, different agents entirely. A background fork that watched one session is the right author for the first kind. It is the wrong author for the second, every time, because it never studied anything. It only watched.
How to talk to the background review
- 01
Name the procedure, not the session.
The nudge fires on a tool-iteration count, not on meaning. A vague instruction lets a judgment call get filed as if it were a workaround.
Instead of
Remember what we just did.
Type this
Save this as a skill: the retry backoff for the webhook that kept timing out. Don't save anything about which voice we used for the post.
- 02
Put a human between the fork and the disk for anything persona-shaped.
write_approval stages every skill_manage write for review instead of committing it straight to disk, which is cheap insurance against a small model misjudging what it learned.
Type this
Set skills.write_approval to true, and show me /skills pending before anything you wrote applies.
Before you let the background review write a skill
- Required:Turn on write_approval if a small model is doing the reviewing.skills.write_approval: true stages every write for a queue you clear by hand.
- Required:Read what it saved, not just that it saved something.A lesson plus the mechanism is useful. A narrated incident is not.
- Required:Keep persona and judgment work out of the automatic loop entirely.Write those by hand, from material you actually studied.
- Required:Check references/ survived if an orchestrator sits in front of Hermes.Paperclip's build strips them; a symlink back to the source fixes it.
- Required:Trust the hub's trust levels over install counts.builtin and official skip the warning panel. community can be forced past a caution, never past a dangerous verdict.
Hermes Agent skills, quick answers
What is a Hermes Agent skill?
A SKILL.md file plus optional references, scripts, templates, and assets that the agent loads on demand. It follows the open agentskills.io standard, the same loading model (metadata always, body on trigger, references on demand) as Claude skills.
What is the Hermes Agent skills hub?
The install and discovery layer behind hermes skills install: official optional skills, GitHub repos and taps (anthropics/skills, openai/skills, huggingface/skills, NVIDIA/skills among the defaults), and community registries including skills.sh, ClawHub, and browse.sh.
Every install runs a security scan first, and trust level decides what a --force flag is allowed to override.
What is the difference between a Hermes Agent plugin and a skill?
A skill is markdown the model reads to learn a procedure. A plugin is Python registered through plugin_loader.py that extends what the runtime itself can do: new tools, request hooks, CLI commands, or memory providers. Skills are optional and loaded by description. Plugins are always active once enabled.
Can Hermes Agent write its own skills?
Yes. A background review nudges every 10 tool iterations by default and can save a workflow as a skill through the skill_manage tool, scoped to a forked agent. A weekly curator marks unused skills stale after 14 days and archives them after 30, and never deletes.
What are the best Hermes Agent skills to use?
The ones whose source sits highest on the hub's trust ladder, official or a verified GitHub tap over an unaudited community install, and, for anything that needs a point of view rather than a procedure, the ones you write yourself from material you actually studied, not the ones a background review saved after watching a single session.
Are Hermes Agent skills compatible with Claude Code skills?
Yes, both implement the open agentskills.io standard. The same SKILL.md module, with the same references and loading map, works in either runtime without a rewrite.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)