Hermes Agent on Telegram: Run Your Agents From a Voice Note
Hermes Agent's Telegram gateway, checked against the v0.21.5 source: pairing and allowlists, chat approvals for risky commands, voice note transcription, and the bot token as a secret worth guarding.
A board is where you decide. A chat is where you operate. Run both through the same agent and the one that never sleeps will swallow the one that is supposed to think, unless you keep them as two different surfaces on purpose.
I run Hermes Agent’s Telegram gateway as my personal assistant. I talk to it by voice, usually away from a keyboard, and it opens issues on my own Paperclip board without me opening the web app. This is what that gateway is actually built from, verified against the source, and the rule I run it by: the chat is where you operate, not where you decide.
Hermes Agent Telegram: what the gateway actually does
Hermes Agent’s Telegram integration is one of 22 platform plugins, built on python-telegram-bot, that turns a Telegram chat into a surface for the same agent that already runs your shell, your memory and your MCP tools. You create the bot through @BotFather, paste the token into ~/.hermes/.env as TELEGRAM_BOT_TOKEN, and the gateway connects by long polling (the default, outbound from your server) or by webhook (inbound, for cloud platforms that need to sleep between messages and auto-wake on traffic).
Nothing about this is Telegram-only plumbing bolted onto a chatbot. The gateway is the same process that talks to more than 25 messaging platforms, and Telegram gets the full feature set: threads and forum topics, streaming message edits, native media, inline keyboards, slash commands and a fallback network transport for when the main connection is unhealthy.
What Telegram touches in a Hermes profile
- ~/.hermes/// HERMES_HOME
- .envsecret// TELEGRAM_BOT_TOKEN, TELEGRAM_ALLOWED_USERS
- config.yaml
- platforms.telegram.extra// mentions, threads, proxy, status
- stt// voice notes in
- tts// voice bubbles out
- mcp_servers// the board, wired as a client
- pairing/// chmod 0600, one row per approved user
Pairing, allowlists and the shared group session
Only two things decide who the bot answers: a numeric Telegram user ID on the allowlist, or a pairing code you approved. TELEGRAM_ALLOWED_USERS is a static, comma-separated list of IDs, read at startup. The more useful path for anyone who is not you is DM pairing: an unapproved sender gets a one-time code back, forwards it to you by any channel, and you run hermes pairing approve telegram <code> without restarting the gateway. Codes expire after an hour, a platform accepts at most three pending codes at a time, five failed approval attempts trigger a one-hour lockout, and the pairing records on disk are chmod 0600. TELEGRAM_ALLOW_ALL_USERS exists and is documented as dev-only, because the bot has shell access.
Groups add a layer most people get wrong once: Telegram’s privacy mode is on by default, so a bot in a group only sees slash commands, replies to its own messages and service events, never ordinary chatter, until you turn privacy off in BotFather and remove and re-add the bot (Telegram caches the old setting on existing memberships). With telegram.require_mention: true, Hermes answers only to a /command, a reply, an @mention or a configured wake-word pattern. Add observe_unmentioned_group_messages: true on top of an allowlisted chat and the bot reads along without replying: unmentioned lines get appended to the shared chat or topic session transcript as context, each one tagged with the sender’s name and ID and carrying its own per-turn instruction that tells the model those lines are context, not commands. A later mention pulls that whole transcript in. That is the closest thing to a group session Hermes has: one transcript per chat or topic, shared by whoever is allowlisted into it.
Approving a risky command from the chat
When Hermes wants to run a command flagged as risky, it posts an inline-keyboard card in the same chat, with four buttons: approve once, approve for the session, always allow, or deny. The card’s header is literal: “Hermes wants to run a command that needs your OK,” with a line underneath stating why it was flagged. Silence is a safe no: the default approval timeout is 300 seconds, and the card tells you so before it expires. This card renders the same way on every messaging surface, Telegram included, through one shared module so the wording never drifts between platforms.
That is the part of “operate from chat” that actually matters. A notification feed tells you what happened. An approval card lets you decide whether it happens, from the same thread you were already in, without switching to a terminal.
Voice notes: speech to text without leaving the chat
A voice note you send on Telegram gets auto-transcribed by whichever speech-to-text provider you configured, then injected into the conversation as text. stt is a top-level section in config.yaml, not a Telegram-specific setting, so the same provider also handles your voice notes on Discord, Signal or WhatsApp:
stt:
enabled: true
provider: "local" # "local" (free) | "groq" | "openai" | "mistral" | "xai"
tts:
provider: "edge" # "edge" (free) | "elevenlabs" | "openai" | ...
Three of them are worth knowing by name: local runs faster-whisper on the machine hosting Hermes and needs no API key, groq calls Groq Whisper and reads GROQ_API_KEY, openai calls OpenAI Whisper and reads VOICE_TOOLS_OPENAI_KEY. Setting stt.enabled: false skips transcription entirely: the gateway still caches the audio file and hands the agent its path instead, useful if you want a custom pipeline (diarization, a different model) rather than Hermes’ own.
Replies come back the same way. TTS output is delivered as a native Telegram voice bubble, the round inline-playable kind, when the provider is OpenAI or ElevenLabs (both produce Opus directly). Edge TTS, the free default, outputs MP3 and needs ffmpeg on the host to convert it to Opus for the bubble; without ffmpeg it still sends, just as a regular audio file.
The bot token is a secret. Treat it like one
TELEGRAM_BOT_TOKEN is marked password: true in the plugin’s own manifest, and for a reason beyond convention: anyone holding it can read every message your bot can see and send messages as it. Hermes ships a dedicated redaction path for this exact string, built into agent/redact.py, that strips the token out of any api.telegram.org/bot<TOKEN>/... URL before it reaches a log line.
That redaction had gaps, and the project’s own regression tests document them: four call sites in the Telegram adapter built their error text straight from the raw exception instead of the redacted version, and the worst of the four persisted the unredacted URL into a dashboard-facing status file, not just a log. A fix closed all four. I bring this up because I have hit the general version of this bug myself: a config.yaml dump inside a Hermes session once wrote my API tokens straight into the session log, and I rotated them. Redaction code is a mitigation, not a guarantee. If a token can print, assume it eventually will, and keep every one of them one /revoke away from dead.
How I run it: the board decides, the chat operates
Three layers, one harness
- 1
Paperclip
The board. Strategic and tactical work: heads run their cycles, specialists deliver against an acceptance bar, heads judge what comes back, and I decide at the gate. This is where I decide.
- 2
Telegram
The chat. Operational work: do this now, approve this command, transcribe this voice note. This is where I operate, mostly away from a keyboard.
- 3
Hermes
The memory. Same harness, same models, same MCP servers, same skills, consulted by both layers above it.
The board is where I decide. The chat is where I operate. The agent’s memory is what both of them read.
The Paperclip board is wired into my Telegram-facing Hermes as an MCP server, the official @paperclipai/mcp-server, which exposes forty-one tools over the board’s API. It sits in the same mcp_servers block of config.yaml that Hermes uses for any other MCP server, with the Paperclip URL and an API key passed in its environment.
Say a line like “write a script covering topic X” into Telegram, by voice or by text, and that sentence becomes a Paperclip issue, through one of those forty-one tools. It is not Hermes doing the work. It is Hermes filing the work where my agent company will pick it up on its own schedule. Operating and deciding stay two different places, on purpose.
Before I built it this way, I explored a different design: a cron job inside Paperclip that would periodically read Hermes’ own session exports and fold them into Paperclip’s memory, a read-model projection instead of a live bridge. It stayed exploration. I never built it. What shipped was simpler and push-based: Telegram talks to the board directly, through one MCP server, the moment I ask.
At my day job, a separate Hermes runs on a cron job instead of a chat: a daily briefing that reads the company’s MCP sources and writes to me, with no chat surface at all. The chat is not the only way to operate an agent. It is the way that lets you interrupt it.
Hermes Agent Telegram vs OpenClaw Telegram
On Telegram specifically, Hermes Agent and OpenClaw both ship it as a first-party channel with real feature depth, but they disagree on who starts the conversation. Hermes only speaks on Telegram when you message it or a cron job fires. OpenClaw’s heartbeat wakes the agent on a timer, by default every 30 minutes (an hour with Anthropic OAuth), and lets it decide whether something is worth a message to you, unprompted.
Where the two differ for a chat-first setup
| Hermes Agent | OpenClaw | |
|---|---|---|
| Behind it | Nous Research, a model lab | OpenClaw Foundation, a 501(c)(3) |
| Wakes on its own | Cron jobs only, no ambient heartbeat | Heartbeat every 30 minutes by default |
| Cost of ambient wake | None: nothing fires without a trigger | About 100k tokens per beat, 2k to 5k with isolatedSession |
| Memory on the chat surface | MEMORY.md, USER.md, full-text session search | MEMORY.md, USER.md, daily notes, dreaming consolidation |
If you want the bot to start the conversation on its own, OpenClaw’s heartbeat does that natively, at a token cost you tune with isolatedSession. If an orchestrator or a cron job is already your heartbeat, as mine is, that feature is pure overhead. The longer comparison, read from both codebases, is in Hermes Agent vs OpenClaw.
Hermes Agent iMessage: same gateway, different bridge
Hermes Agent reaches iMessage through BlueBubbles, a separate open-source macOS server that bridges iMessage to any device. It is the same agent, the same memory, the same approval cards, behind a different bridge with a hard requirement Telegram does not have: an always-on Mac, signed into Messages.app, running BlueBubbles Server. You point Hermes at it with a server URL and password, the same way you point it at a Telegram bot token, and the gateway treats it as one more platform adapter.
If your operating surface has to be a phone number your family already texts, this is the honest reason to accept the Mac dependency. If it does not, Telegram asks for a free bot and a VPS, nothing more.
Hermes Agent WhatsApp: the bridge that can get you banned
Hermes Agent’s default WhatsApp integration is Baileys, an unofficial bridge that emulates a WhatsApp Web session. No Meta developer account, no business verification, and a real (if small) risk that WhatsApp flags the number. Hermes documents two modes: a dedicated bot number (lower risk, cleaner for multiple users) or your own number in a self-chat (fastest to set up, worst for anything beyond testing). For anyone running this as more than a personal experiment, Hermes also ships the official path, the WhatsApp Business Cloud API, which trades the setup speed for a Meta Business account, a public webhook, and no ban risk.
Telegram has no equivalent split. There is one bot, one token, one official API, which is part of why it is the surface I run.
Before Telegram becomes your operating surface
- Required:You have a board somewhere else that already decides.Telegram is for operating, not for planning. If you only have a chat, you have a todo list with extra steps.
- Required:You are on an allowlist or DM pairing, not TELEGRAM_ALLOW_ALL_USERS.A bot with shell access and no allowlist is a public terminal with a friendly name.
- Required:You have read the approval card's four choices before you need them.Once, session, always, deny. Knowing the difference before a 2am approval request beats reading it then.
- Required:You picked an STT provider on purpose.Local needs no key and runs on your own machine. Groq and OpenAI need an API key and bill per use. Decide before the first voice note, not after.
- Required:The bot token lives in .env, not in a message you asked the agent to read aloud.Redaction exists and it has had gaps. Rotate on any suspicion, the same day.
Hermes Agent on Telegram, quick answers
How do I connect Hermes Agent to Telegram?
Create a bot through @BotFather, put the token in TELEGRAM_BOT_TOKEN inside ~/.hermes/.env, add your numeric Telegram user ID to TELEGRAM_ALLOWED_USERS, and start the gateway. python-telegram-bot handles the connection, by long polling unless you configure a webhook.
Is OpenClaw better than Hermes Agent on Telegram?
They differ on who starts the conversation, not on channel depth. OpenClaw's heartbeat can message you first, on a timer. Hermes only speaks on Telegram when you message it or a cron job fires. Pick the one that matches whether you want an ambient assistant or a reactive one.
Does Hermes Agent support voice messages on Telegram?
Yes. Incoming voice notes are transcribed by the stt provider in config.yaml (local, free; groq, needs GROQ_API_KEY; openai, needs VOICE_TOOLS_OPENAI_KEY) and injected as text.
Replies can come back as native voice bubbles through tts, using OpenAI, ElevenLabs or the free Edge TTS provider (the last one needs ffmpeg installed to produce the round bubble instead of a plain audio file).
Does Hermes Agent work with WhatsApp or iMessage too?
Yes, as separate platform adapters with the same memory and approval system behind them. iMessage goes through BlueBubbles and needs an always-on Mac. WhatsApp's default path is an unofficial bridge with a small ban risk; a WhatsApp Business Cloud API path exists for anyone who needs the official route instead.
Is my Telegram bot token safe with Hermes Agent?
Hermes marks it as a secret and has dedicated code to strip it from log lines and error text. That redaction has had documented gaps in the past, since fixed. Treat the token like any other credential: keep it out of .env files anyone else can read, and rotate it the moment you suspect it printed anywhere.
What happens if I don't approve a risky command on Telegram?
Nothing runs. The approval card times out after five minutes by default, and the card itself says so: silence is a safe no, not a default yes.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)