What Is an Agent Harness? Everything in the Agent That Isn't the Model
An agent harness is all the code around an AI model that turns it into an agent: the loop, the tools, the context, and the controls. What each part does, real harnesses side by side, and why the harness decides more than the model.
A model can think. It cannot open a file, run a test, or remember yesterday. Everything that lets it do those things is the harness.
An agent harness is all the code around an AI model that turns it into an agent. It runs the loop that keeps calling the model, hands the model tools to act with, decides what the model sees on every turn, and enforces what the model is allowed to do. The model produces text. The harness is what lets that text touch your repo.
Claude Code is a harness. So are Codex CLI, OpenCode, pi, Kiro, and Hermes Agent. When people say “my agent got better,” most of the time the harness changed and the model stayed the same.
I run harnesses every day. I shipped a 13-app crypto fintech in 70 days, solo, with agents in the loop. I built a seven-agent pipeline in OpenCode on open models, wrote my own model-handoff package for pi, and keep a Hermes agent running on a VPS. This page is the plain definition I wish I had found first.
The word comes from horses
A horse is strong. Left alone in a field, that strength moves nothing you care about. Put a harness on it, hitch it to a cart, give someone the reins, and the same strength delivers a load to a place you chose.
The model is the horse. It brings the raw capability: reading code, reasoning about it, writing more of it. The harness is the collar, the straps, the cart, and the reins. It points the capability at a job, carries the result somewhere useful, and keeps a human hand on the direction. Swap in a stronger horse with a broken harness and you still go nowhere.
Agent equals model plus harness
LangChain put the definition in four words. It is the cleanest version anyone has published, and it is the one I use in harness engineering.
The definition
The test at the end of that box is the useful part. You cannot touch the weights. You can touch every other line. That is why the harness is where your leverage lives.
The loop is the smallest part
Open a real harness and the first surprise is how small the agent loop is. Strip any coding agent down and you find the same few lines.
while the model asks for tools:
run the tools it asked for
append the results to the context
call the model again
A 2026 study that took Claude Code apart, Dive into Claude Code, says it plainly: the core is “a simple while-loop that calls the model, runs tools, and repeats.” Most of the code lives around that loop: a permission system with seven modes and a classifier, a five-layer compaction pipeline for context, four extension mechanisms (MCP, plugins, skills, hooks), subagent orchestration, and session storage.
So when someone says they built an agent in fifty lines, they built the loop. The harness is the other ninety-something percent, and it is the part that decides whether the agent ships work or ships confident garbage.
The four parts of every harness
Every serious harness has the same four parts. The names change between products. The jobs do not.
What a harness is made of
- 01
The loop
Calls the model, runs the tools it asked for, feeds back the results, and repeats until the work is done. It also decides when to stop: a step cap, a tool the model calls to finish, or a check that passes.
pi: no step cap, runs until the model stops calling tools OpenCode: a steps budget per agent
- 02
The tools
The model's hands. Read a file, edit a file, run a shell command, call an API. Every tool definition costs tokens on every turn, so fewer and sharper beats more.
pi ships four: read, write, edit, bash Claude Code ships many, plus MCP servers
- 03
The context
What the model sees each turn: the system prompt, the instruction files, the history, the tool output. Plus memory, the files that survive when the session ends. This is where most tokens die.
AGENTS.md / CLAUDE.md loaded every session compaction when the window fills
- 04
The controls
What the agent may do without asking. Permissions per tool, approval prompts, hooks that run before or after an action, and the gate a human must open before the next phase.
edit: ask, bash: allow for git, deny for rm a review gate before merge
If you want the deep version of the third part, context engineering covers what the model should see, and AGENTS.md covers the memory file. The fourth part is where spec-driven development earns its keep: the spec is the contract the gate checks against.
You already use one
Every coding agent you have heard of is a harness wrapped around someone’s model. They differ in which of the four parts they let you change.
Six harnesses, one definition
| Harness | Models | What stands out in the harness |
|---|---|---|
| Claude Code (Anthropic) | Claude, or any Anthropic-compatible endpoint | Permission modes, hooks, skills, subagents, plugins, MCP |
| Codex CLI (OpenAI) | OpenAI models | Open source, sandboxed execution, approval modes, AGENTS.md |
| OpenCode | 75+ providers, local models | Agents with their own model, temperature, steps and permissions |
| pi | Any provider, bring your own key | Four tools, a ~150-word system prompt, sessions as JSONL |
| Kiro (AWS) | Claude, GPT and open-weight models | Spec workflow with approval gates, steering files, hooks |
| Hermes Agent (Nous Research) | Pluggable model backends | Always-on, persistent memory, writes its own skills |
Is Claude Code a harness?
Yes. Claude Code is Anthropic’s harness for Anthropic’s models, and that pairing is the point of the product. The loop, the permission system, the compaction pipeline, the skills and hooks: all harness. The model on the other end is Claude.
That pairing has a commercial edge too. A Claude plan covers usage inside Anthropic’s own harnesses (Claude Code, Cowork, the apps). Since April 4, 2026, a third-party harness that logs in with your plan bills as extra usage instead. So “which harness” and “which subscription” became the same decision. I wrote that trade up in is Claude Max worth it.
The fastest way to feel the harness is to keep the model fixed and change everything else. Run the same Claude model in Claude Code and in pi. Same weights, different system prompt, different tools, different context rules. You get a different agent.
Harness, framework, model
Three words get mixed up constantly. They are three layers.
Framework
- 01A kit for building agents: LangChain, LangGraph, the Claude Agent SDK
- 02You write code against it
- 03Ships building blocks, not a working agent
Harness
- 01A working agent runtime: Claude Code, OpenCode, pi, Kiro
- 02You configure it and point it at work
- 03Ships the loop, tools, context and controls, assembled
The model sits under both. You can build a harness with a framework, and some harnesses are built that way. When someone searches for a “LangChain agent harness,” they usually mean the harness they will assemble from LangChain’s parts.
Why the harness matters more than the next model
Every lab ships a better model every few months, and everyone gets it the same day. The harness is the part that is yours, and it compounds with your judgment.
I saw it most clearly in my own OpenCode setup. I gave review to an open model at temperature zero, with edits denied and a checklist of eight categories. It could only answer APPROVE or REQUEST_CHANGES, so it could not quietly fix what it was supposed to report. Planning went to a different model with the shell turned off. None of those rules came from the models. All of them came from the harness, and the same models without them would have been a different agent. The full config is in OpenCode vs Claude Code.
That is also why “what is the best AI for coding” is the wrong first question. The better one is which harness, and then which model in each phase of the work. I answer it in the best AI for coding.
Agent harness, quick answers
What is an agent harness?
An agent harness is all the code around an AI model that turns it into an agent: the loop that keeps calling the model, the tools it acts through, the context it sees on each turn, the memory that survives the session, and the controls that decide what it may do. Agent = Model + Harness.
Is Claude Code a harness?
Yes. Claude Code is Anthropic's agent harness, built around Claude models: the loop, permission modes, compaction, skills, hooks, subagents, plugins and MCP support are all harness.
A Claude Pro or Max plan covers usage inside Anthropic's own harnesses. Third-party harnesses that log in with the plan bill as extra usage since April 4, 2026.
What is the difference between an agent harness and an agent framework?
A framework (LangChain, LangGraph, the Claude Agent SDK) is a kit you write code against to build an agent. A harness is a working agent runtime you configure and run, such as Claude Code, OpenCode, pi or Kiro. You can build a harness with a framework.
What is the difference between the model and the harness?
The model is the trained network that reads and writes text. The harness is everything else: prompts, tools, the loop, context management, memory files, permissions and hooks. You cannot change the model's weights, but you can change every part of the harness.
What is harness engineering?
Harness engineering is the practice of improving everything around the model instead of waiting for a better model: defending the context window, keeping tools few and sharp, giving the agent durable memory, and making the loop verify its own work. The full guide is in the harness engineering article.
Can I use any model in any harness?
Technically, often yes. Open harnesses like OpenCode and pi accept many providers, and Claude Code can point at Anthropic-compatible endpoints such as Ollama.
Commercially, no. Subscriptions decide it: a Claude plan covers Anthropic's harnesses, ChatGPT Plus and Pro can be used in participating third-party tools since September 29, 2026, and an API key works anywhere at per-token prices.
What is the pi agent harness?
pi is an open-source coding agent harness by earendil-works with four tools (read, write, edit, bash), a system prompt of about 150 words, no MCP by default, and sessions saved as JSONL. It runs against any provider with your own key, which makes it the easiest harness to read end to end.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)