Skip to content
← articles
updated PaperclipHermes AgentAI AgentsMulti-Agent SystemsAgent Orchestration

I Built a Company of AI Agents in Paperclip. Here Is What Broke.

Thirteen agents on Paperclip and Hermes Agent: a CEO, a chief of staff, five heads and the specialists they route work to. What Paperclip is, how the org is wired, and every failure it took before the loop closed on its own.

My content team is MrBeast, Gary Vaynerchuk and Neil Patel. My offers go through Alex Hormozi. Russell Brunson builds the funnels, and Richard Feynman reviews anything that tries to teach. None of them has ever heard of me. They are agents, thirteen of them, running on Hermes Agent inside Paperclip, on a VPS I pay for by the month.

I have walked dozens of engineers and engineering leaders through this org. They all ask the same three questions. What is Paperclip, actually. Does it work. What breaks. This is the long answer, from the week I installed it in June to July 22, the first day the loop closed without me pushing it.

The org, in numbers

13
agents1 CEO, 1 chief of staff, 5 heads, 6 specialists
~6
weeksfrom install to a loop that closes alone
2
npm packagesadapters I had to write myself
1
human gatepublishing, by design
The org chart took an afternoon. Everything else on this page took the six weeks.

What is Paperclip AI?

Paperclip is an open-source, MIT-licensed app for running AI agents as an organization. You get an org chart with reporting lines, issues that agents assign to each other, monthly budgets, approvals, an audit trail, and a heartbeat that wakes each agent on a timer or the moment work lands on it. The repo calls it “the open-source app everyone uses to manage agents at work,” and it sits at about 95k stars on GitHub.

Here is the part the landing pages skip. Paperclip does not think. It has no agent loop of its own. Every agent connects through an adapter to a runtime that does the actual work: Claude Code, Codex, OpenCode, OpenClaw, a plain HTTP endpoint, or, in my case, Hermes Agent.

That split decides most of your problems. When something goes wrong, the first question is always which side of the adapter it happened on. Most of my bugs lived exactly on the adapter.

Why a developer ends up running a company of agents

I did not start here. I climbed to it, and every rung removed the friction the one below it left behind.

The ladder that ended at an agent company

  1. 1

    Vibe coding

    "Copy from ChatGPT, paste into the editor, hope."

    The fear goes away. The context never arrives.

  2. 2

    A real product with agents

    "A 13-app crypto fintech in 70 days, specified first, built by agents."

    New development anchored on AI. Every session still starts blank.

  3. 3

    A second brain

    "An LLM wiki plus skills that capture what each session learns."

    Context survives. Execution still waits for me.

  4. 4

    Persona skills

    "An expert I studied, packaged as a skill the model can load."

    Precision. I am still the one driving every session.

  5. 5

    Autonomous agents

    "Hermes plus Paperclip: context plus MCP equals execution without babysitting."

    The work moves while I am not there.

The ladder that ended at an agent company: 5 progressive levels, from the most basic (Vibe coding) to the most advanced (Autonomous agents).

The ladder is cumulative. The Paperclip org reads the second brain, runs on the persona skills, and reaches the world through MCP servers. Skip a rung and the top one has nothing to stand on. The fintech case study is rung two in detail.

The reason I needed rung five is boring. My job changed from developer to one-person business. One person who needs a blog, a YouTube channel, offers and funnels, on a fixed number of hours a week, with a family that gets the rest of them. Agents that wait for me to type do not solve that. Agents that wake up on their own do.

The org: one CEO, a chief of staff, five heads, six specialists

The org chart, as it runs in Paperclip

  • CEO agent// escalates to me, the human, when something is above its pay grade
    • Gwynne Shotwellchief of staff// turns five heads into one digest
      • MrBeasthead// Library: content and attention
      • Sal Khanhead// Academy: teaching
      • Marc Louhead// Arsenal: self-serve products
      • DHHhead// Build: done-for-you work
      • Pieter Levelshead// Ventures: my own products
    • Gary Vaynerchukspecialist// content
      • Arthur Miller// narrative review
      • Richard Feynman// didactic review
    • Neil Patelspecialist// editorial intelligence and SEO
    • Alex Hormozispecialist// offers
    • Russell Brunsonspecialist// funnels

Three roles, three different jobs.

Heads own an outcome tied to a number, one head per line of business. They never do the craft. They pick the bet, open an issue for a specialist with the acceptance bar written inline and a deadline, and judge what comes back. If a head is writing copy or pulling raw data, its bar was not sharp enough.

Specialists own one craft. Each one carries a persona skill and its own MCP servers. A specialist has a single function: think, and consult data through MCP.

The chief of staff collects the approved wins from the five heads and hands the CEO one decision-ready digest. That role exists because of a failure. My first CEO agent tried to read the whole board in one run and timed out holding everything.

The heads carry real names too, in their SOUL.md. The same persona routing that sharpens a specialist works one layer up. A head called MrBeast judges a content bet differently from a head called “Content Manager,” and that difference is the point.

Heads judge, specialists think

The split between the two is the most important design decision in the org. It is also the one most multi-agent demos skip, because they give one agent the work and the grade.

Head

  1. 01Owns an outcome tied to a number
  2. 02Starts cycles on its own timer
  3. 03Writes the acceptance bar into the issue
  4. 04Accepts the return or sends it back
  5. 05Never does the craft

Specialist

  1. 01Owns one craft
  2. 02Wakes only when work is assigned to it
  3. 03Reads real data through its MCP servers
  4. 04Delivers against the bar
  5. 05Never judges whether its work passes
The person who did the work never grades it.

The specialist does not judge whether the work passes. Whoever knows what they wanted judges it.

Three things fall out of that rule. Confirmation bias drops, because the agent that produced the work is not the one approving it. Each context stays narrow, one craft per specialist. And narrow contexts run fine on smaller, cheaper models, which matters when thirteen agents wake up on a schedule.

Same harness, two purposes

Every agent in this setup runs on the same harness. Same models, same context, same MCP servers, same skills. The only thing that changes is the purpose the agent answers to.

Paperclip is the strategic and tactical layer: the heads’ weekly cycles, where work gets planned, routed and judged. Hermes on Telegram is the operational layer: “do this for me now.” I talk to it on my phone, usually by voice, and it has the Paperclip board wired in as an MCP server. When something cannot wait for a head’s cycle, I say “a script on topic X” into Telegram and it opens the issue on the board for me.

That is the whole mental model. Paperclip is where I decide. Telegram is where I operate. Hermes is the memory both of them consult.

The four files every agent reads

Skills route an agent by topic. These four files define the job. Every agent in the org gets the same bundle shape, injected before each wake.

One agent's instruction bundle

  • agents/<agent>/instructions/
    • SOUL.mdwho// persona, organizing principle, posture, anti-patterns
    • AGENTS.mdwhere// place in the org, delegation, gates, safety
    • HEARTBEAT.mdloop// what to do on every wake, step by step
    • TOOLS.mdhow// Paperclip API, skills, scoreboard, who to route to

The built-in Hermes adapter ignored all four. Paperclip ships a hermes_local adapter that spawns the Hermes CLI with the task as the prompt and parses what comes out. It never sent the instruction bundle. My agents woke up with a task and no identity.

So I wrote the adapter I needed and published it: @felipefontoura/paperclip-adapter-hermes-local-plus, a drop-in replacement that registers under the same adapter type. It prepends the bundle on every wake, keeps each skill’s references/ folder (Paperclip’s runtime build stripped everything except SKILL.md), reads the wake context from the object Paperclip actually fills, recovers runs that came back blank, and reports real token counts and cost on the run page.

It also fixed a leak nobody would have guessed. Hermes auto-discovers an AGENTS.md in its working directory, and the agents were running inside Paperclip’s own app folder. Every wake was injecting Paperclip’s internal AGENTS.md into the prompt. Giving each agent its own scratch directory took input from 9,044 tokens to 6,992. That is 22 percent of every wake, gone.

I tried patching the built-in adapter first. The patches piled up with no version and nowhere to share them, which is why I now own the seam instead. I also built the other shape: Hermes behind HTTP, through a gateway adapter and a small skills bridge that serves Paperclip’s skills to Hermes over HTTP, with no Hermes CLI inside the Paperclip container. The org runs on the local adapter today.

What broke, in order

Nothing on this list is exotic. Each one is the kind of failure you only meet when a fleet runs for days instead of a demo running for a minute.

Six weeks of an agent company

  1. First runProduction worked. Publication did not.

    Neil and Gary ran about fifty heartbeats and delivered a synthesis, a channel baseline and three complete video packages. Every package stopped at the one step that needs a human: publishing to the channel. The bottleneck moved from me as producer to me as approver.

  2. First runAn agent escalated over a silent board

    The human board went quiet for about five days. Gary noticed he had only pinged the board and never escalated through the chain, so he opened an issue for the CEO agent. The CEO agent authorized a stopgap and rescoped the week. That one issue unblocked the cycle.

  3. First runTwo agents did the same work

    Neil created three issues three minutes after the board had created the same three, which left six issues for three videos. Later a parallel instance of Gary ran the same heartbeat and produced the deliverable twice. He caught his own duplicate and marked it obsolete.

  4. RebuildThe org was wiped and the doc did not notice

    A rebuild of the adapter and the server stack reset the database. My notes said five chiefs in production. Postgres said one agent, the CEO, idle. The foundation working is not the same as the org existing.

  5. TuningThe same model scored up to 26 percent lower

    On my own offer-design shootout, GLM-5.1 scored 9.55 through OpenCode and 7.0 to 7.9 through Paperclip plus Hermes. Same model, same persona. The runtime had sent it two identities.

  6. TuningFour minutes of silence killed an agent

    Liveness was inferred from stdout. A long silent synthesis looked dead, the orchestrator spawned a duplicate, the duplicate stole the lock, and the real run's write-back bounced. A fifteen-line keepalive fixed it.

  7. Tuning157 tools, zero calls

    The agents swore their tools did not exist. The execution log said the MCP backends were cold-spawning past a sub-second discovery window. A caching pooler in front of the gateway fixed it.

  8. Org v2A head re-delegated forever

    When a specialist handed finished work back, the head saw a sensor task assigned to itself and delegated it again. In a loop. The fix was one rule: an in-review issue assigned to you is a return. Judge it and advance, never re-delegate.

  9. July 22The loop closed on its own

    MrBeast oriented the week. Neil read the sensors and returned a ranking of eight angles. MrBeast accepted and advanced to Gary. Gary produced the hooks. A separate report went up to the chief of staff and MrBeast closed the parent. Zero loops, zero orphans, a shortlist waiting at the human gate.

Most of these were adapter and orchestration bugs. The models were rarely the problem.

Three of these deserve more than a line.

The runtime sent the model two identities

The drop looked like a model problem. It was a plumbing problem. I pulled the exact prompt out of Hermes’ session database. The system role was a 14k-character prompt that opens with Hermes introducing itself as Hermes Agent. The persona, “You are Alex Hormozi, Chief Offer Architect,” arrived later as a user message. The persona never appeared in the system prompt at all.

Models are trained to weigh the system role. So the model got “helpful generalist” from the system and “aggressive closer” from the user, and it picked the safe average. Correct, generic, and missing everything that made the persona worth loading. The fix was to give each agent one identity, in one place, loaded where the runtime treats it as identity.

The lesson I took from it: testing models without testing where your prompt lands is measuring half of the system.

The bottleneck moved to me

An agent company does not remove you. It moves you from producer to approver. My agents produced three video packages and then waited at the only gate that needs a human, which was me, and I was not there. If you do not redesign the approval step, the org produces into a queue. This is the same dip teams hit when AI moves the bottleneck to review.

I kept the gate anyway. Agents produce. I publish. A shortlist waiting for a yes is a much better problem than an agent posting to my channel on its own.

The agent is an unreliable narrator

When the tools went missing, the agents told me a detailed, confident story about upstreams that never connected. I believed them for an hour. The execution log told a different story: a cold-spawn race. The same thing happened with the agent that died in silence. Its own report said nothing useful. The heartbeat code did.

The rule I run on now: never debug the agent’s account of a failure. Read the log. Both fixes went public: the pooler is mcp-pooler, and the keepalive is an open pull request against Nous Research’s adapter.

The loop that finally closed

Paperclip wakes an agent in three ways: when an issue is assigned to it, when something it depends on resolves, and on a timer. The design that works uses all three with one discipline. Only heads start cycles on a timer. Specialists never start anything. They wake when work is assigned to them and nowhere else.

One cycle of the Library head

Input

MrBeast wakes on his timer

  1. ORIENTRead the engine, pick the lens

    Read the goal, the open issues and the last ranking. Pick which sense to look through this week: self, competitor, trend, frontier, community or client.

  2. ROUTEOpen an issue for a specialist

    Assign it, write the acceptance bar and the deadline inline, and create it open. Work born closed wakes nobody.

  3. WORKThe specialist reads the world

    Neil reads the sensors through his MCP servers and delivers. Then he sets the issue to in review and hands it back to the head, which wakes the head.

  4. JUDGEAccept and advance, or send back

    A return is judged, never re-delegated. Accepted work becomes the next demand for the next specialist in the chain, here Gary.

  5. REPORTReport up, close the parent

    The head files a new report issue for the chief of staff and closes the cycle's parent issue itself. The parent never leaves the head's hands.

Output

A decision-ready shortlist at the only human gate: publish

One cycle of the Library head: flow of 5 steps from “MrBeast wakes on his timer” resulting in “A decision-ready shortlist at the only human gate: publish”.

Each step in that pipeline is a scar. “Create it open” exists because work born done woke an assignee that had nothing to do. “Hand it back to the head” exists because Paperclip only lets the assignee change an issue, so a specialist’s return never woke anyone. “Never re-delegated” is the loop. “The parent never leaves the head’s hands” is the time a head passed the whole cycle up to the chief of staff halfway through.

Where it stands today

The loop closed on July 22. The board has been quiet since that afternoon: the last heartbeat anywhere in the org is from that day. I took on a bigger role in my day job, and this org existed to make a one-person business run on the hours I had left. When those hours went elsewhere, the org went quiet with them. That is the honest property of an agent company with one human gate. No human at the gate, no company.

Only one of the five heads, MrBeast, has ever run a cycle, and that is on purpose. Library was the engine with real demand behind it. The other four manage lines of business I have not validated yet, a SaaS among them. Waking a head to run a business that does not exist yet is theater. Each one goes live when its business does, over the next few months.

I am turning it back on as my YouTube channels and social accounts get active again. This time the org gets a narrower job: run a creator’s operation, the daily cadence of channels and posts, with me back at the gate every morning.

When Paperclip is worth it

Before you build an agent company

  • Required:
    You have more than one front that needs a different judgment.Content, offers and funnels are three judgments. One person with one task does not need a hierarchy. It only adds latency.
  • Required:
    Two or three persona skills already work in a plain session.The org multiplies skills you have. It does not create them.
  • Required:
    Your context lives in files the agents can read.A second brain, an AGENTS.md, a canon of what you never do. Knowledge that repeats belongs there, not inside an agent.
  • Required:
    Every specialist reaches real data through MCP.An agent without a sensor is blind and guessing. The MCP server is the bridge from the sensor to the data.
  • Required:
    You know exactly where the human gate is, and you show up to it.Mine is publishing. If you are not there every morning, the org produces into a queue.
  • Required:
    You start with the smallest loop that closes.One input you type by hand, one agent that thinks and proposes, you decide, it ships. Automate data collection last, not first.
Two or more unchecked and you are building an org chart, not a company.

Paperclip, quick answers

What is Paperclip AI?

Paperclip is an open-source, MIT-licensed app for running AI agents as an organization: an org chart, issues agents assign to each other, budgets, approvals and heartbeats that wake agents on a timer or when work arrives. It does not run a model itself. Each agent connects through an adapter to a runtime such as Claude Code, Codex, OpenClaw or Hermes Agent.

Does Paperclip run the agents?

No. Paperclip schedules, routes and records. The thinking happens in the runtime behind each agent's adapter.

That is why most production bugs live on the adapter: identity, wake context, liveness and token accounting all cross that boundary.

What is the difference between Paperclip and Hermes Agent?

Paperclip is the org layer. Hermes Agent is an agent runtime with its own tools, memory, skills and messaging gateways.

In my setup every Paperclip agent runs on Hermes, and a separate Hermes on Telegram is my personal assistant, wired to the Paperclip board so I can open issues by voice.

Is Paperclip a zero-human company?

Mine is not, by design. Agents produce, route and judge each other's work, but publishing goes through one human gate. A shortlist waiting for a yes is a better problem than an agent posting to your channel unsupervised.

Which model should Paperclip agents use?

Let the runtime own the model, not the orchestrator. My adapter stopped passing a model and reads it from Hermes' own config. Over these weeks the fleet ran on GLM-5.1 through Z.AI and later DeepSeek V4 Flash through OpenRouter.

Narrow specialists run fine on smaller models. Test where your persona lands in the prompt before you blame the model: the same model scored up to 26 percent lower in my setup because the runtime handed it two identities.