The Best Model for Hermes Agent Is the Cheapest One That Survives Your Runtime
There is no best model for Hermes Agent, only a best model for the job routed through the runtime you actually run. My own shootout shows the same model losing two points for no reason except where the persona landed.
Every “best model for X” article is the same article with the brand name swapped. Mine is not going to be that, because I ran a test that shows the question is broken. The same model, same task, same judge, scored 9.55 in one runtime and as low as 7.0 in another. Nothing about the model changed between those two runs.
The question was never which model is smartest. It was which model survives the harness you actually run it through.
I run Hermes Agent in production. Here is the shootout that taught me the lesson, what I actually ran in the months since, and the method I use instead of a leaderboard.
Hermes Agent model: the runtime owns it, not the orchestrator
The model for a Hermes instance lives in one place: the model block in ~/.hermes/config.yaml, under HERMES_HOME. Nothing else should be allowed to set it. In June my orchestrator passed a model and a reasoning-effort flag into every Hermes run. Its model management was buggy, and the effort flag was not even one the Hermes chat command accepted, so it killed the wake before the agent did anything. My adapter stopped passing either and let Hermes read both from its own config.
The model belongs to the runtime’s config file, not to whatever orchestrates it. Paperclip, cron, a Telegram message, none of them should carry a model argument into Hermes. They start the work. Hermes decides what runs it, because Hermes is the thing that has to live with the consequences: the context budget, the reasoning effort, the tool-calling format. Get this backwards and every other decision in this article is noise.
Best model for Hermes Agent: my shootout, and why the scores moved
On June 8, I ran my own judged offer-design task, the same prompt and the same rubric, across four model-and-runtime combinations. I was trying to answer “which model is best.” What I got was a lesson about where the persona lands.
One task, one judge, four combinations
| Model, via which runtime | Score out of 10 |
|---|---|
| GLM-5.1, via OpenCode | 9.55 |
| Opus, via the Claude Code adapter | 8.88 |
| GLM-5.1, via Paperclip plus Hermes | 7.0 to 7.9 |
| gpt-5.5, via Paperclip plus Hermes | 6.7 |
The pattern is the finding: it is not that GLM is better than gpt-5.5, or that Opus is the strongest of the four. It is that routing GLM-5.1 through Hermes cost it between 1.6 and 2.5 points against the score the same model got through OpenCode. I pulled the actual prompt out of Hermes’ session database to find out why. The system prompt opened with 13,985 characters introducing the agent as Hermes Agent. The persona I wanted, “You are Alex Hormozi, Chief Offer Architect,” arrived as a 10,831-character plain user message. The model got two identities and picked the safe average: correct, generic, and missing everything the persona was supposed to add.
That root cause has nothing to do with model intelligence. A smarter model does not fix a persona sent to the wrong slot. The fix is structural: identity goes in SOUL.md, where Hermes treats it as identity, not in a user message competing with the system prompt for the model’s attention.
What I actually ran, month to month
Here is the honest log, not a recommendation list. In June, the fleet ran on GLM-5.1 through Z.AI, the month I moved model selection out of the orchestrator and into config.yaml for good. By July 21, I had switched the runtime to deepseek/deepseek-v4-flash through OpenRouter. Also in June, gpt-5.4 through Codex, riding a ChatGPT Plus subscription, was one of the sources of tokens for the same agents.
None of those three is “the best model.” They were the best model for that fleet, at that cost point, at that moment, which is a different claim and the only honest one to make about a project that reconfigures itself every few weeks.
My own opinion, from using both: GLM handled long, multi-source synthesis with more depth than DeepSeek. DeepSeek was perfectly fine for simpler execution work, the kind where the plan is already made and the model just needs to carry it out without losing the thread. That split, deep synthesis versus precise execution, is the same split the model-handoff piece wires into pi: different phases of work deserve different models.
Route by role, not by fleet
The lever I trust more than a smarter model is not picking one model for the whole fleet in the first place. In my Paperclip org, heads own an outcome and judge; specialists own one craft and execute inside a narrow, fixed context. A specialist with a tight brief and a persona skill carrying its judgment does not need the most expensive model available. The judgment is already written down. What is left for the model to do is narrower, and narrow work is exactly what a cheap model does well.
Hermes gives you a built-in version of “ask more than one model and let a stronger model judge.” Its Mixture-of-Agents mode, /moa, sends one question to multiple reference models and hands their answers to a separate aggregator model that synthesizes the final response. The shipped default preset pairs gpt-5.5 through the OpenAI Codex provider and deepseek/deepseek-v4-pro through OpenRouter as the two references, with anthropic/claude-opus-4.8 through OpenRouter as the aggregator. All three are overridable, and the shape is the lesson even if you never touch the default: cheaper models generate candidates, one capable model makes the final call. That is role-based routing, built into the runtime, one config block away from a pattern I had to build by hand at the org level.
Hermes Agent local model: Ollama, vLLM and llama.cpp through the custom provider
Running Hermes against a model on your own hardware, “ollama hermes agent” in the searches that bring people to this page, goes through the same model block, with provider set to custom. Hermes’ custom provider profile exists specifically for this: any OpenAI-compatible endpoint, with Ollama, vLLM and llama.cpp as the named cases, plus aliases in the provider registry (ollama, local, vllm, llamacpp) that all resolve to the same profile. The shape of the config is this:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"
No API key is required for a local Ollama endpoint. The provider profile fills in a sensible default reasoning effort when you leave it unset, and it widens the clamp on reasoning-effort values to the full OpenAI-compatible set, because a custom endpoint’s own vocabulary is not something Hermes can discover ahead of time.
The one filter that matters more than raw model quality here is tool calling. Hermes is agentic: it edits files, runs commands, calls MCP servers. A local model that cannot reliably emit tool calls can chat with you and do nothing else, no matter how good its prose is. Shop for that capability first, then for everything else. If your box is slow, check the timeouts before you blame the model: Hermes’ stale-call detector is auto-disabled for local endpoints, and you can set a per-provider request timeout with providers.<id>.request_timeout_seconds.
Job, model class, why: a method instead of a ranking
How I actually decide, not a leaderboard
| Job | Model class | Why |
|---|---|---|
| Narrow specialist, one craft, judgment already in SOUL.md | Cheapest model that passes the test | The context is narrow and the judgment is fixed. More model does not buy more quality here. |
| Long synthesis across many sources | A model with patient, deep reasoning (GLM-tier, in my experience) | Thin reasoning shows up fastest on long synthesis, not on short execution steps. |
| Aggregating several agents' output, including /moa's aggregator role | Your most capable, most expensive model | One call per cycle, not per turn, so the premium is paid once. |
| Local-only, cost-sensitive, tool-driven work | Whatever your hardware runs with reliable tool calling | A model that cannot call tools cannot drive Hermes at all. Chat quality is the wrong axis to shop on. |
| Anything you have not profiled yet | The cheapest model that passes your own test, run through the real runtime | A model that scores well in a plain chat window can lose two points the moment it is routed through your harness. |
Best models for OpenClaw: the same method, the same traps
The question comes up for OpenClaw too, and the answer does not change shape. OpenClaw treats model providers as extensions, with Ollama, LM Studio and vLLM in that list alongside OpenRouter, plus subscription OAuth for ChatGPT, Codex and the Claude CLI, and a model failover feature that switches providers automatically when your primary one errors out.
Nothing about that list makes one model universally best. The same two filters apply: verify tool-calling support before you shop on chat quality, especially for a local model, and test through the actual runtime, not a plain chat window, because OpenClaw’s heartbeat and context settings can shift a model’s effective performance the same way Hermes’ system prompt shifted mine.
How to run your own shootout
Replace the leaderboard with a test you control
- Required:Test through the real runtime, never a plain chat window.The harness can cost a model two points before it writes a word, exactly as it did to GLM-5.1 under Hermes with the persona in the wrong slot.
- Required:Judge with a fixed, written rubric before you see outputs.My shootout used one judged task, scored the same way, across every model and runtime combination. Without that discipline the comparison is just vibes.
- Required:Separate model quality from runtime cost.A model that scores a point lower but costs a tenth as much is still the right call for a narrow specialist doing fixed work.
- Required:Confirm tool-calling support before shopping on chat quality.Especially for local models. A model that can only chat cannot drive an agentic loop, no matter how fluent it reads.
- Optional:Re-run the shootout when the runtime version changes, not only when the model does.Hermes ships close to twice a week. A score from June is a snapshot of that commit, not a permanent verdict.
Best model for Hermes Agent, quick answers
What is the best model for Hermes Agent?
There is no single best model. The best model is the cheapest one that passes a test run through your actual Hermes configuration, not a plain chat window. My own shootout showed the same model losing up to 2.5 points purely from where Hermes placed the persona in the prompt.
What model does Hermes Agent use by default?
None is forced on you. Hermes reads the model from the model block in ~/.hermes/config.yaml, and it supports 38 provider plugins plus any OpenAI-compatible endpoint through the custom provider. Set it there, not in whatever orchestrator or gateway wakes the agent.
Can I run Hermes Agent with Ollama?
Yes. Set provider to custom and base_url to your Ollama endpoint, typically http://localhost:11434/v1, with no API key required. The provider registry also accepts ollama, local, vllm and llamacpp as aliases for the same profile.
What is the best local model for Hermes Agent?
Whichever one your hardware can run with reliable tool calling. Hermes is agentic: it edits files and runs commands through tool calls, and a local model without solid tool support can only chat. Verify that capability before comparing anything else.
What are the best models for OpenClaw?
The same method applies as for Hermes. OpenClaw treats providers as extensions, including Ollama, LM Studio and vLLM alongside OpenRouter and subscription OAuth, plus model failover if one provider errors out. Test through OpenClaw's real runtime before trusting a model's reputation from elsewhere.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)