Writing
How to Set Up Hermes Agent: Hooks, Memory, and Skills
updated 2026-08-12
I spent two weeks moving from Claude Code to Hermes Agent as my primary agent setup. I use it to run a private knowledge vault, a task system, scheduled briefs on Telegram, and client work on personal accounts in off hours. This is what I learned, written down so you do not have to learn it the hard way.
This is not a features page. Hermes's own docs are the source of truth for commands and APIs. This field guide is the operational layer: the pitfalls that cost me hours, and the settings that actually matter.
For the August 2026 free-model snapshot, see Free AI Models for Hermes Agent.
What this guide covers
- Install and key paths
- Provider setup (subscription vs API)
- Memory, soul.md, and skills
- Hooks: nudge vs enforcement
- Profiles, cron, and delegation
- Desktop output control
- What I would do differently
Official docs stay canonical. If a command here disagrees with the docs, trust the docs.
What Is Hermes Agent?
Hermes Agent is an open-source AI agent framework by Nous Research.
It runs in your terminal, a native desktop app, messaging platforms
(Telegram, Discord, Slack, WhatsApp, iMessage, Signal, Matrix, Teams,
Email), and IDEs. It works with any LLM provider — OpenRouter,
Anthropic, OpenAI, Google, DeepSeek, xAI, local models, and 20+ others.
What makes it different from Claude Code or Codex CLI:
- Self-improving through skills — Hermes saves reusable procedures
as skills that load into future sessions. It learns from experience.
- Persistent memory across sessions — remembers who you are, your
preferences, environment details, and lessons learned. No more
starting from zero every session.
- Multi-platform gateway — the same agent runs on every messaging
platform with full tool access, not just chat.
- Provider-agnostic — swap models and providers mid-workflow;
credential pools rotate across multiple API keys automatically.
- Profiles — run multiple independent Hermes instances with
isolated configs, sessions, skills, and memory.
- Hooks — a full lifecycle hook system with two enforcement tiers:
pre_llm_call nudges the model by injecting context, transform_llm_output
enforces by modifying the response after the model finishes.
Installation
# Shell installer — sets up uv, Python, the venv, and the launcher
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
After installation:
hermes setup --portal # Fastest path if you have a Nous subscription
hermes setup # Full wizard — pick model + provider
hermes model # Change model/provider later
hermes doctor # Health check
hermes status # Current configuration status
Key Paths
~/.hermes/config.yaml Main configuration (settings — never secrets)
~/.hermes/.env API keys and secrets ONLY
~/.hermes/skills/ Installed skills
~/.hermes/soul.md Agent personality file
~/.hermes/agent-hooks/ Shell hook scripts
~/.hermes/hooks/ Gateway event hooks (Python)
~/.hermes/plugins/ Desktop app plugins
~/.hermes/state.db Session store (SQLite + FTS5)
~/.hermes/sessions/ Session transcripts
~/.hermes/logs/ Gateway and error logs
~/.hermes/auth.json OAuth tokens and credential pools
If you use profiles (recommended for separating personal projects from client instances), each profile lives under ~/.hermes/profiles/<name>/ with the same layout.
Provider Setup: The Subscription vs. API Trap
This is the #1 gotcha. A ChatGPT Plus or Claude Pro web plan does not
automatically give Hermes an API.
| Provider | Web plan | What Hermes can use |
|---|---|---|
| Nous Portal | Subscription includes API | Same account. Fastest path: hermes setup --portal |
| OpenAI | ChatGPT Plus/Pro is not API | Codex OAuth (Codex models) or a key from platform.openai.com |
| Anthropic | Claude Pro is not API | Claude Max with extra credits via OAuth, or a key from console.anthropic.com |
| xAI / Grok | SuperGrok is not API | A key from console.x.ai |
| Google / Gemini | Gemini Advanced is not API | AI Studio or Vertex |
| OpenRouter | N/A | One key, many models |
| Ollama | Local | Local weights. No token bill |
The Nous Portal exception: If you have a Nous subscription, Hermes
works out of the box. The subscription is the API. Run hermes setup --portal.
For everything else: Confirm you have an API path, not only a web login.
Provider docs change. Check the current Hermes quickstart before you assume a plan includes OAuth.
OpenRouter: The Path of Least Resistance
If you want access to multiple models (GPT, Claude, Grok, Llama, etc.)
without creating separate API accounts at each provider, use
OpenRouter. One account, one API key, pay-per-token, ~$10-20 deposit.
# 1. Go to openrouter.ai, create an account, deposit $10-20
# 2. Generate an API key (starts with sk-or-v1-...)
# 3. Add to Hermes
hermes auth add openrouter --type api-key --api-key <key> --label "OpenRouter"
# 4. Models are accessed as openrouter/<provider>/<model>
# e.g. openrouter/anthropic/claude-sonnet-4.6
# e.g. openrouter/openai/gpt-4o
# e.g. openrouter/xai/grok-3
# 5. Or use the interactive picker
hermes model
OpenRouter also has free-tier models (marked with :free in the model
ID). Listings change. Do not copy an old id into a production loop.
See Free AI Models for Hermes Agent for an August 2026 selector snapshot, then verify the live listing.
Test from the CLI with whatever :free id is live:
hermes chat -q "What is 2+2?" --provider openrouter -m "PROVIDER/MODEL:free"
Local Models (Ollama)
If you have a capable machine, run models locally for zero ongoing cost:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | bash
# Pull a model
ollama pull qwen3:14b
# Create a dedicated Hermes profile for local use
hermes profile create local-scribe
# Configure the profile to use Ollama (set provider=custom, base_url=http://localhost:11434/v1)
hermes --profile local-scribe config set model.provider custom
hermes --profile local-scribe config set model.base_url http://localhost:11434/v1
hermes --profile local-scribe config set model.default qwen3:14b
On a 32GB Apple Silicon desktop, a 14B local model can take minutes per query with Hermes's full system prompt. Slow for interactive use. Fine for batch work. Zero tokens. Zero cost.
Configuration
Edit with hermes config set section.key value or hermes config edit
for the full file. Never hand-edit config.yaml — a stray indent can
corrupt the file and break the live gateway.
Most-Used Config Sections
| Section | Key Options |
|---|---|
model | default, provider, base_url, api_key, context_length, aliases |
agent | max_turns (90), tool_use_enforcement, verify_on_stop |
terminal | backend (local/docker/ssh/modal), cwd, timeout (180) |
compression | enabled, threshold (0.50), target_ratio (0.20) |
display | skin, interface (cli/tui), language, show_reasoning, show_cost |
approvals | mode (smart/manual/off), timeout, cron_mode |
stt | enabled, provider (local/groq/openai/mistral) |
tts | provider (edge/elevenlabs/openai/minimax/mistral/gemini) |
memory | memory_enabled, user_profile_enabled, provider, write_approval |
security | redact_secrets, tirith_enabled, website_blocklist |
delegation | model, provider, max_concurrent_children, max_spawn_depth |
hooks | Shell hook configuration (see below) |
Memory: Persistent Context Across Sessions
Memory is injected into every turn — no re-reading a 200-line config
file to recover context. Split into two targets:
- user: Who you are, role, technical level, communication style,
hard rules, recurring preferences.
- memory: Environment facts, tool quirks, conventions, lessons.
Keep memory compact and high-signal. It's injected every turn, so
bloated memory costs tokens every single message. Don't store task
progress or completed-work logs — use session_search for those.
Soul.md: Agent Personality
Write a soul.md file at ~/.hermes/soul.md that defines your agent's
personality. Mine is direct, adversarial, minimalist, and explicitly
told to push back on bad ideas. This is your agent's character — make
it match how you actually want to work.
Skills: Procedural Memory
Skills are reusable procedures that load when relevant. They're
Hermes's version of Claude Code's .claude/skills/ — but better,
because they have categories, trigger descriptions, and can carryreferences/, templates/, and scripts/ directories.
Skills live in ~/.hermes/skills/<category>/<name>/SKILL.md.
---
name: my-skill
description: "Use when <trigger>. <one-line behavior>."
version: 1.0.0
author: Your Name
license: MIT
platforms: [linux, macos, windows]
---
# My Skill
## When to Use
...
## Steps
1. ...
## Pitfalls
1. ...
The first 57 characters of the description appear in the system prompt
as a trigger hint. Keep the trigger self-contained in that window.
The Skills Hub
Hermes ships with hundreds of skills — built-in, optional, and community sources (Claw Hub, LobeHub, gstack, Anthropic, OpenAI, HuggingFace, NVIDIA). Counts change. Check the Skills Hub for the live catalog. The full catalog lives at the
Skills Hub, browsable
by category (Apple, AI Agents, Creative, GitHub, Media, MLOps, Productivity,
Research, Smart Home, Social Media, Software Dev) and platform.
The built-in skills are the ones worth studying first. They are maintained by the Hermes team and show the SKILL.md format at its best. Optional skills cover a wide range, from Notion API integration to Manim video generation. You install them the same way: drop the directory in ~/.hermes/skills/<category>/<name>/.
For agents and power users, the key insight is that skills are
procedural memory that compounds. Every time you solve a hard problem
or discover a non-obvious workflow, saving it as a skill means the next
session starts with that knowledge loaded. This is the core thesis behind
using Hermes as a second-brain operating system — your procedures become
reusable artifacts, not session-specific ephemera.
The Surfaces
Hermes runs on multiple surfaces — the same agent core drives all of
them:
- Desktop app (
hermes desktop) — native Electron app, streaming
chat, session list, Cmd+K palette, drag-and-drop files.
- CLI (
hermes) — interactive terminal chat. - Ink TUI (
hermes --tui) — terminal UI with docked widgets. - Web dashboard (
hermes dashboard) — admin panel with messaging
channels, MCP catalog, webhooks, memory, profile builder.
- Gateway — connects to Telegram, Discord, Slack, WhatsApp,
iMessage, Signal, Matrix, Teams, Email, and more.
- OpenAI-compatible proxy (
hermes proxy) — point Codex CLI,
Aider, or any script at it. No API key needed.
Telegram Gateway Setup
The gateway lets Hermes run on messaging platforms with full tool
access. Telegram is the most common setup:
# Interactive (if you have a terminal):
hermes gateway setup
# Non-interactive (from an agent context, no TTY):
# 1. Add bot token to .env
sed -i.bak 's/^# TELEGRAM_BOT_TOKEN=.*/TELEGRAM_BOT_TOKEN=<token>/' ~/.hermes/.env
# 2. Install as persistent service
hermes gateway install
# 3. Verify
hermes gateway status
Once the gateway is running, Hermes responds to your Telegram messages
with full tool access — terminal, file operations, web search, code
execution, everything. It's not a chatbot; it's your agent living on
your phone.
Hooks: Nudge vs. Enforcement
This is the feature that changed everything for me — but not the way
I first thought. The distinction between nudging the model and
enforcing behavior is the most important lesson in this entire guide,
and it cost me four failed sessions to learn it.
I had a standing protocol (filler-word tracking) that my agent kept
skipping under the pressure of real work. Three sessions in a row.
Memory didn't fix it. Skills didn't fix it. Both rely on the model
choosing to comply.
So I built a pre_llm_call shell hook — a reminder injected into
every user message before the model sees it. The hook fired
correctly. The model received the reminder. **And then ignored it
again.** Four sessions of failure. The hook was working; the model
just chose not to comply. A pre_llm_call hook is a Post-it note
on the model's monitor — it can still walk past it.
The fix was transform_llm_output — a hook that fires after the
model produces its final response but before the user sees it.
It can append to or replace the response using classical code. The
model literally cannot skip it because it runs after the model is
done. No model cooperation required.
The Nudge-vs-Enforcement Distinction
| Approach | Layer | How It Works | Can Model Skip It? |
|---|---|---|---|
| Memory | System prompt | Text injected every turn | ✅ Yes — background noise |
| Skills | Conditional load | Model loads when relevant | ✅ Yes — model decides not to load |
| System prompt rules | Static instructions | Model reads at start | ✅ Yes — prioritizes immediate task |
pre_llm_call hook | Pre-response | Injects context before model responds | ✅ Yes — it's a nudge, not a gate |
transform_llm_output | Post-response | Modifies response after model finishes | ❌ No — runs after the model is done |
The rule: if you need the model to remember to do something every
response, pre_llm_call will help but won't guarantee it. If you need
a specific output format or footer to appear every response
regardless of model behavior, use transform_llm_output as a Python plugin. A shell hook for this event was broken as of July 2026 (issue #67890). Check current docs before you depend on the shell path.
The architecture that works: nudge with pre_llm_call (cheap,
no tokens), enforce with transform_llm_output (deterministic, code
not model). Together they're reliable. This is a sellable pattern for
client deployments — "the agent's behavior is mechanically enforced,
not just suggested."
The Full Hook List
Hermes has four hook systems. Here are all of them:
Shell Hooks (config.yaml — CLI + Gateway)
These fire in both CLI and gateway sessions. Declare in~/.hermes/config.yaml, point at shell scripts. Any language (Bash,
Python, Go binary — anything with a shebang).
| Hook | Fires When | Can Do What |
|---|---|---|
pre_tool_call | Before any tool executes | Block the tool call |
post_tool_call | After any tool returns | Observe/log |
pre_llm_call | Before model responds (once per turn) | Inject context into user message |
post_llm_call | After model finishes responding | Observe/log |
pre_verify | When agent edited code, before finish | Keep agent going (run checks) |
on_session_start | New session created | Observe |
on_session_end | Session ends | Observe |
on_session_finalize | Session torn down | Observe |
on_session_reset | /new or /reset used | Observe |
subagent_start | delegate_task child spawned | Observe |
subagent_stop | Child agent finishes | Observe |
pre_gateway_dispatch | Gateway receives msg, before auth | Skip/rewrite/allow msg |
pre_approval_request | Before approval prompt | Observe |
post_approval_response | After approval decision | Observe |
transform_tool_result | After tool returns, before model sees | Rewrite result |
transform_terminal_output | Inside terminal, pre-truncation | Rewrite output |
transform_llm_output | After model finishes, before delivery | Rewrite response |
Gateway-Only Hooks (fire only in Telegram/Discord/Slack/etc)
| Hook | Fires When |
|---|---|
gateway:startup | Gateway process starts |
session:start | New messaging session |
session:end | Session ends |
session:compress | Context compression completed |
agent:start | Agent begins processing |
agent:step | Each iteration of tool loop |
agent:end | Agent finishes processing |
reaction:added | Emoji reaction added |
reaction:removed | Emoji reaction removed |
command:* | Any slash command |
Practical Hook Recipes
Recipe 1: Protocol Nudge (pre_llm_call)
Inject a reminder into every user message before the model sees it.
Important: this is a nudge, not enforcement — the model can still
ignore it. I built one of these to remind my agent to include a filler-word
tracking report in every response. The hook fired correctly every turn.
The model ignored it four sessions in a row. Use this for gentle reminders;
use transform_llm_output (Recipe 1b) when you need the output to actually
appear.
#!/usr/bin/env python3
"""Protocol nudge — injected into every user message via pre_llm_call."""
import sys
import json
REMINDER = """════════════════════════════════════════════════════════════════
HARD RULE REMINDER — DO NOT SKIP
1. Always run the filler tracking scan on every response.
2. Always update Tasks.md when completing or cancelling work.
3. Always check git log before reporting task status.
These apply to EVERY response. No exceptions.
════════════════════════════════════════════════════════════════"""
def main():
try:
json.loads(sys.stdin.read() or "{}") # consume stdin
except Exception:
pass
print(json.dumps({"context": REMINDER}))
if __name__ == "__main__":
main()
Register in config.yaml:
hooks:
pre_llm_call:
- command: /usr/bin/env python3 /absolute/path/to/hook.py
timeout: 5
Allowlist for non-TTY sessions in ~/.hermes/shell-hooks-allowlist.json:
{
"approvals": [
{
"event": "pre_llm_call",
"command": "/usr/bin/env python3 /absolute/path/to/hook.py"
}
]
}
Test: hermes hooks test pre_llm_call
Recipe 1b: Protocol Enforcement (transform_llm_output plugin)
This is the one that actually works when the model doesn't cooperate.transform_llm_output fires after the model finishes but before the
user sees the response — it can append to or replace the response using
classical code. The model cannot skip it.
Important: transform_llm_output as a shell hook is currently broken
(GitHub issue #67890 — as of July 2026 the shell-hook parser dropped transform stdout). Use a Python plugin instead, then re-check the docs. A later Hermes release may have fixed the shell path.
Plugin structure:
~/.hermes/plugins/my-enforcer/
├── plugin.yaml
└── __init__.py
plugin.yaml:
name: my-enforcer
version: "1.0.0"
description: "Append a formatted report to every response"
author: Your Name
provides_hooks:
- transform_llm_output
- pre_llm_call
__init__.py (simplified — see below for the key pattern):
"""Enforcer plugin — appends a formatted footer to every response."""
# Module-level stash: pre_llm_call stores the user message here,
# transform_llm_output reads it later (transform_llm_output doesn't
# receive user_message as a kwarg, so we need to pass it through).
_CURRENT_USER_MESSAGE = ""
def pre_llm_call(user_message: str, session_id: str, **kwargs) -> None:
"""Stash the user message for transform_llm_output."""
global _CURRENT_USER_MESSAGE
_CURRENT_USER_MESSAGE = user_message or ""
def transform_llm_output(response_text: str, session_id: str,
model: str, platform: str, **kwargs) -> str:
"""Append the formatted report to the response.
This fires AFTER the model finishes but BEFORE the user sees it.
The model cannot skip this — it's classical code running after
the model is done.
"""
# Build your report from _CURRENT_USER_MESSAGE
report = build_report(_CURRENT_USER_MESSAGE)
return response_text + report
def register(ctx):
ctx.register_hook("transform_llm_output", transform_llm_output)
ctx.register_hook("pre_llm_call", pre_llm_call)
Enable: hermes plugins enable my-enforcer
⚠️ Gotcha: Plugin changes require a full app restart. The Hermes
desktop app is a single long-running Python process. Python caches
imported modules for the life of the process — starting a new session
(or saying the wake word) does NOT re-import plugins. If you edit__init__.pyafter the app is running, your changes won't take effect
until you completely quit the app (Cmd+Q on macOS) and relaunch.
Clear__pycache__first to force a clean compile:rm -rf ~/.hermes/plugins/my-enforcer/__pycache__
>
This is the #1 plugin debugging time-sink. The plugin shows as
"enabled," the code is correct, it works in testing — but the app
doesn't fire it. The running process has a stale module. Full restart,
not a new session.
The model can ignore every pre_llm_call nudge, skip every
protocol, forget every standing rule — and the report still appears,
because it's appended by code, not by the model.
Recipe 2: Block Dangerous Commands (pre_tool_call)
Reject destructive terminal commands before they run:
hooks:
pre_tool_call:
- matcher: "terminal"
command: /absolute/path/to/block-rm-rf.sh
timeout: 5
#!/usr/bin/env bash
payload="$(cat -)"
cmd=$(echo "$payload" | jq -r '.tool_input.command // empty')
if echo "$cmd" | grep -qE 'rm[[:space:]]+-rf?[[:space:]]+/'; then
printf '{"action": "block", "message": "blocked: rm -rf / is not permitted"}\n'
else
printf '{}\n'
fi
Recipe 3: Auto-Format After Every Write (post_tool_call)
hooks:
post_tool_call:
- matcher: "write_file|patch"
command: /absolute/path/to/auto-format.sh
#!/usr/bin/env bash
payload="$(cat -)"
path=$(echo "$payload" | jq -r '.tool_input.path // empty')
[[ "$path" == *.py ]] && command -v black >/dev/null && black "$path" 2>/dev/null
printf '{}\n'
Recipe 4: Audit Every Response (post_llm_call)
Log every response to an external system — client visibility into what
the agent did:
#!/usr/bin/env python3
import sys, json, httpx
def main():
payload = json.loads(sys.stdin.read() or "{}")
extra = payload.get("extra", {})
response = extra.get("assistant_response", "")
session_id = payload.get("session_id", "")
# Log to external API, Slack, dashboard, etc.
try:
httpx.post("https://your-logging-service.com/api", json={
"session_id": session_id,
"response": response[:500],
}, timeout=5)
except Exception:
pass
print("{}")
if __name__ == "__main__":
main()
Recipe 5: Output Sanitizer (transform_llm_output)
Strip PII or enforce house style before the response reaches the user:
#!/usr/bin/env python3
import sys, json, re
def main():
payload = json.loads(sys.stdin.read() or "{}")
response = payload.get("extra", {}).get("response_text", "")
# Strip email addresses
cleaned = re.sub(r'\S+@\S+', "[EMAIL REDACTED]", response)
print(json.dumps({"response": cleaned}))
if __name__ == "__main__":
main()
Recipe 6: Startup Checklist (gateway:startup)
Run a checklist every time the gateway boots — check overnight cron
failures, summarize logs, ping you if anything broke:
Create ~/.hermes/BOOT.md with natural-language instructions, then
create a gateway hook that fires on gateway:startup and runs a
one-shot agent to execute the checklist. If nothing's wrong, the agent
replies with [SILENT] and you hear nothing.
Hook Isolation: No Cross-Tool Leaks
Critical safety property: Hermes hooks only fire for Hermes. They're
declared in ~/.hermes/config.yaml and loaded only by the Hermes
runtime. Other tools — Cursor, Claude Code, Codex — have their own
separate hook systems and never read Hermes config.
I learned this the hard way with Claude Code. The Stop hook in.claude/settings.json fired globally, in every repo Claude Code
touched — including when Cursor was working in an adjacent directory.
Cursor hit the hook, looped on it, and burned a session. With Hermes,
this can't happen. The hook system is completely isolated.
Hook Setup Checklist
- Write the hook script (Python, Bash, anything with a shebang)
chmod +xthe script- Register in
~/.hermes/config.yamlunderhooks: - Use absolute paths —
~is NOT expanded - Pre-allowlist in
~/.hermes/shell-hooks-allowlist.jsonfor non-TTY - Test:
hermes hooks test <event> - Verify:
hermes hooks list - Restart Hermes — but know the difference. Shell hooks (config.yaml
changes) take effect on the next session. Python plugins do not —
the desktop app caches imported modules for the life of the process.
A new session or wake word reuses the old module. To pick up plugin
code changes, fully quit the app (Cmd+Q) and relaunch. Clear
__pycache__ first: rm -rf ~/.hermes/plugins/<name>/__pycache__
Profiles: Multiple Independent Instances
Run separate Hermes instances with isolated configs, sessions, skills,
memory, and API keys:
hermes profile create studio # Personal projects
hermes profile create client1 # Per-client profile
hermes --profile studio # Run with a specific profile
Each profile has its own config.yaml, .env, skills, memory, and
sessions. Different API keys = different rate-limit windows = parallel
work. This is how you deploy Hermes for multiple clients without
cross-contamination.
Cron Jobs: Scheduled Autonomy
Hermes cron jobs run the agent itself on a schedule — not just a script.
Full tool access, multi-platform delivery.
# Create via the cronjob tool (available in-session)
cronjob(
action='create',
name='Daily Task Reminders',
schedule='0 7 * * *', # 7am daily
deliver='telegram', # CRITICAL: set explicitly
prompt='You are Hermes Agent. Read Tasks.md...',
enabled_toolsets=['file', 'terminal'] # limit tools to reduce cost
)
Critical pitfall: The default deliver is origin (the current
session). If you don't explicitly set deliver='telegram', the job
runs, does its work, and the result vanishes into a local file you
never see. Always set deliver explicitly.
Cost note: On a subscription (Nous Portal), agent sessions cost
tokens. A daily brief cron job that runs the agent for 5 minutes every
morning burns tokens every day. For mechanical tasks (no reasoning
needed), a Python script + launchd/cron is the zero-token alternative.
Use agents for reasoning, scripts for mechanics.
Delegation: Parallel Subagents
Hermes can spawn subagents in isolated contexts — each gets its own
conversation, terminal session, and toolset. Only the final summary
returns to the parent.
# Single task
delegate_task(goal='Research the competition and write a summary', context='...')
# Parallel batch (up to 3)
delegate_task(tasks=[
{'goal': 'Audit the backend code', 'context': '...'},
{'goal': 'Write API documentation', 'context': '...'},
{'goal': 'Run the test suite and report failures', 'context': '...'},
])
Use delegation for: reasoning-heavy subtasks, work that would flood
your context with intermediate data, independent parallel workstreams.
Don't use delegation for: a single tool call, mechanical multi-step
work (use execute_code), or tasks needing user interaction.
Importing From Another Agent
If you're coming from Claude Code or Codex CLI, Hermes has a built-in
import:
hermes import-agent
This imports config, credentials, and some skill structures
automatically. Run it first, then do manual mapping for what it
doesn't cover.
The migration has five layers:
- Operating instructions (CLAUDE.md/AGENTS.md) → memory + skills +
cron + stays in repo file
- Skills (.claude/skills/) → ~/.hermes/skills/ with upgraded structure
- Scheduled automation (launchd/cron) → Hermes cron jobs
- Memory (session amnesia) → Hermes persistent memory
- Multi-account → Hermes profiles
The test for each line of your CLAUDE.md: "Would removing this cause a
real mistake?" If memory or a skill handles it, cut it. If it's about
the vault structure itself, keep it. Typical result: 200+ lines shrinks
to 40-60 lines.
The Daily Brief
I run a hybrid architecture: a Python script gathers data (tasks, calendar, weather) and delivers via Telegram at 7am. Zero tokens. The
agent-based cron jobs are paused to conserve subscription credits.
For a full agent-based daily brief (smarter synthesis, multi-platform
delivery), use a Hermes cron job. For a mechanical data-gathering
brief, a Python script is the zero-cost alternative.
Voice Integration
STT (Voice → Text)
Voice messages from messaging platforms are auto-transcribed. Local
faster-whisper is free and private:
stt:
enabled: true
provider: local
local:
model: base # tiny, base, small, medium, large-v3
TTS (Text → Speech)
Edge TTS is free and requires no API key:
tts:
provider: edge
OpenAI, ElevenLabs, MiniMax, Mistral, and Gemini are also supported
(some require API keys).
Controlling Desktop Output: Styling Agent Responses with CSS
The Hermes desktop app renders agent responses as markdown — but it also
intercepts certain link formats and renders them as rich UI widgets. If
you're deploying Hermes for clients or building a polished second-brain
workflow, you'll eventually want to control how those widgets look. Here's
what I learned the hard way.
The Preview Card Widget
When you include a #preview/file:///absolute/path link in a response,
the desktop app's MarkdownLink component intercepts it and renders aPreviewAttachment card — a styled box with an icon, the filename, and
an "Open preview" button. This is great for giving users one-click access
to files, but the widget derives its label from the file path, not
from the markdown link text. So [Click me](#preview/file:///foo/Bar.md)
renders as a card labeled "Bar" (the filename extracted from the path),
not "Click me" (the link text you wrote).
The card's filename span has CSS classes min-w-0 flex-1 truncate — it
grows to fill available space and truncates long names. Tasks.md fits cleanly. A longer name such as Weekly Review.md truncates.
The Fix: Desktop Plugin CSS Injection
The desktop app supports disk plugins — plain ESM JavaScript files
dropped in ~/.hermes/desktop-plugins/<id>/plugin.js. These are loaded
at runtime with full DOM access, including the ability to inject <style>
tags. No build step, no repo clone.
Here's a minimal plugin that hides the filename in preview cards so only
the icon and button remain:
// ~/.hermes/desktop-plugins/clean-preview-cards/plugin.js
const CSS = `
/* Hide the filename span in preview-attachment cards.
The card container has max-w-160; the filename span has
truncate + flex-1. */
div[class*="max-w-160"] > span.truncate.flex-1 {
display: none !important;
}
/* Shrink the card to fit just the icon + button. */
div[class*="max-w-160"] {
max-width: fit-content !important;
width: auto !important;
}
`
export default {
id: 'clean-preview-cards',
name: 'Clean Preview Cards',
register(ctx) {
const style = document.createElement('style')
style.setAttribute('data-plugin', 'clean-preview-cards')
style.textContent = CSS
document.head.appendChild(style)
return () => style.remove()
}
}
Save it. The app watches desktop-plugins/, loads the file within seconds,
and hot-reloads on every save. If it doesn't appear, run ⌘K →
Reload desktop plugins.
How I Found the Right Selector
I couldn't modify the app's source (it's a packaged build). No CDP port
for live DOM inspection. So I traced the rendering chain through the
source code:
- Plugin appends
[·](#preview/file:///...)markdown link MarkdownLinkcomponent (markdown-text.tsx:259) intercepts
#preview/ hrefs via previewTargetFromMarkdownHref()
- Renders
<PreviewAttachment target={target} />
(preview-attachment.tsx)
previewName(target)extracts the filename from the path- Renders: icon span + filename span (
min-w-0 flex-1 truncate) + button
The selector div[class*="max-w-160"] > span.truncate.flex-1 is unique —max-w-160 appears only once in the compiled CSS, only in the
PreviewAttachment component. The !important is necessary because
Tailwind's utility classes have high specificity.
The Broader Lesson: Three Layers of Output Control
When you need an output to appear every time and look a specific way,
there are three layers, each with different enforcement:
- Prompt/soul instructions — the model should include it. Fragile.
Works until it doesn't. The model can forget, skip, or rephrase.
transform_llm_outputplugin — code appends the content after
the model finishes. The model can't skip it. But the content is still
markdown — the desktop app renders it, and you're at the mercy of its
widget components.
- Desktop plugin CSS — code modifies the DOM after the app renders.
Full control over visual appearance. The strongest layer for styling.
For my filler-protocol footer (a table with filler stats + three preview
cards), I use all three: the soul says to include the report, thetransform_llm_output plugin appends it mechanically, and a desktop
plugin injects CSS to hide the filenames in the cards. Each layer is a
backstop for the one above it.
Key Pitfalls
- The link text is ignored.
[·](#preview/...)and
[Click here](#preview/...) produce identical cards. The label comes
from the file path, not the markdown text.
- URL-encode paths with spaces. The remark/CommonMark parser
terminates URLs at the first unencoded space. Use quote(path, safe="/")
in Python. This ate the middle link when I had three links on one line.
- Desktop plugins hot-reload; Python plugins don't. A desktop plugin
file save triggers a reload within seconds. A Python plugin
(__init__.py) change requires a full app restart (Cmd+Q), because the
desktop app caches imported modules in sys.modules for the process
lifetime. Clear __pycache__ first.
- No CDP in packaged builds. The DevTools Protocol port (9222) only
opens in dev-server mode. For inspecting the running app's DOM, either
launch an isolated dev instance or trace the source code.
What I'd Do Differently
- Start with
transform_llm_output, notpre_llm_call. This is
the hardest-won lesson in this guide. I had a standing protocol the
agent needed to follow every response. I tried memory — failed.
Tried skills — failed. Built a pre_llm_call hook that injected a
reminder every turn — the hook worked, the model ignored it. Four
sessions of failure. The problem wasn't the hook, it was the layer:
pre_llm_call is a nudge (the model can still skip it), not
enforcement (the output is modified by code after the model
finishes). The fix was a transform_llm_output plugin that
appends the formatted report to every response mechanically. If
you need an output to appear every time, don't rely on the model
remembering to include it — append it with code. See Recipe 1b
above. This is the single most important architecture pattern in
this guide for anyone deploying agents for clients: **nudge with
pre_llm_call, enforce with transform_llm_output.**
- Use OpenRouter from day one. I wasted time trying to connect
subscriptions (ChatGPT, Grok) that don't include API access.
OpenRouter is $10-20, one key, every model. Start there.
- Keep memory lean. Memory is injected every turn. Bloated memory
costs tokens every single message. Store facts and preferences,
not task progress or completed-work logs.
- Write a real soul.md. The personality file shapes every
interaction. Mine is direct, adversarial, and explicitly told to
push back on bad ideas. If you want an honest operator, not a
yes-man, say so in the soul.
- Use profiles for clients. Do not mix personal projects and client work in the same Hermes instance. Separate profiles mean separate memory, sessions, skills, and API keys. All of this runs on personal accounts.
- Test hooks before deploying.
hermes hooks test <event>fires
the hook against a synthetic payload. Run it before restarting.
Malformed JSON is silently ignored — a broken hook is worse than
no hook because you think it's working.
- Know the restart rules. There are two kinds of hooks and they
have different restart requirements. Shell hooks (config.yaml
changes) take effect on the next session. Python plugins do not —
the desktop app caches imported modules for the life of the process.
I spent an entire session debugging a plugin that was "enabled" and
worked in testing but didn't fire in the app. The code was right. The
app just had a stale module from before my edits. **A new session is
not a restart.** Quit the app (Cmd+Q), clear __pycache__, relaunch.
This is the first thing to check when a plugin works in testing but
not in the app.
transform_llm_outputrequires streaming OFF. This is the
single biggest gotcha with response-transform plugins. When
display.streaming: true (the default), the CLI and desktop app
stream the response token-by-token as the model generates it. The
transform_llm_output hook fires after the model finishes — but
by then the original (unmodified) text has already been displayed.
The CLI sees response_previewed=True and skips printing the
transformed version. Result: the hook fires (proven by log
files), the return value is correct, but the user never sees it.
Fix: hermes config set display.streaming false. The response
will appear as a single block after the model finishes (slightly
less "live" feel), but transform_llm_output modifications will
appear every time. This is a Hermes limitation, not a plugin bug —
but it should be documented prominently because it's not obvious
and it silently breaks every response-transform plugin.
Getting Support from Hermes / Nous
If something breaks or doesn't work as documented, you'll need to file a support request. Here's how.
Where to Get Help
Discord is the primary support channel. Hermes is open-source and community-driven; the Nous Research team hangs out in Discord and responds to real issues.
- Discord server: discord.gg/nous-research (look for #hermes or #support)
- GitHub issues: github.com/nousresearch/hermes-agent (for reproducible bugs with code diffs)
- Docs: hermes-agent.nousresearch.com/docs (always check here first — most "bugs" are just config)
Before You Post: The Checklist
Hermes Discord has guidelines (posted in #rules or #support-guidelines). Read them. You'll need:
- Hermes version —
hermes --version - OS and platform — macOS/Linux/Windows, desktop app/CLI/Telegram gateway
- Relevant config — (never paste API keys, but DO paste
config.yamlsections, hook declarations, plugin.yaml, etc.) - Error message or log output — where exactly did it fail? Paste from
~/.hermes/logs/agent.log - Steps to reproduce — exact commands/interactions that triggered the problem
- What you expected vs. what happened — the gap
How to Write a Good Support Request
Bad (vague, low signal):
Hermes isn't working. My plugin doesn't fire. Help?
Good (clear, reproducible):
Platform: macOS desktop app, version 0.12.3
Issue: Plugin hook `transform_llm_output` declared in plugin.yaml and registered in register(ctx), shows as enabled in `hermes plugins list`, but never invoked at runtime.
Steps to reproduce:
1. Place a plugin with `transform_llm_output` hook in ~/.hermes/plugins/
2. Ensure plugin.yaml declares it: `provides_hooks: [transform_llm_output]`
3. Verify registration: `register(ctx)` calls `ctx.register_hook("transform_llm_output", fn)`
4. Start a new session, send a message
5. Hook never fires (confirmed by: response unchanged, no logs, manual import works)
What I expected: Hook fires after LLM responds, modifies response before delivery
What actually happened: Response delivers unchanged, no hook invocation
Logs: [paste relevant section from ~/.hermes/logs/agent.log]
What NOT to Do
- ❌ Don't ask for help with custom code issues (their problem, not Hermes's) unless it's a Hermes API question
- ❌ Don't post in random channels — look for #hermes, #support, or #plugins first
- ❌ Don't expect instant replies — Nous team is lean, volunteer-driven, working async across timezones
- ❌ Don't paste API keys, tokens, or full config with secrets (redact them)
- ❌ Don't report "it doesn't work" without reproduction steps (they can't help)
Common Issues & Self-Fixes
Plugin changes aren't taking effect:
- Clear the plugin cache:
rm -rf ~/.hermes/plugins/<name>/__pycache__ - Restart the app fully (Cmd+Q on macOS, not just a new session) — the Python process caches imported modules for its lifetime
Config changes aren't working:
- Are you editing
~/.hermes/config.yamldirectly? ✓ Changes take effect next session - If you used
hermes config set, did you test withhermes config get <key>? - Check YAML with
hermes doctor. Do not hand-editconfig.yamlif you can usehermes config set.
Hook isn't firing:
- Is it listed in
hermes hooks listorhermes plugins list? - Check logs:
tail -100 ~/.hermes/logs/agent.log | grep -i "hook\|error" - Can you import the hook function directly?
python3 -c "from your_module import hook_fn"✓ = registration issue, ❌ = code issue
Memory/skills not loading:
- Memory: injected into every turn, check with
hermes memory show - Skills: only loaded when relevant. Test manually:
hermes skill load <name>
Memory vs. Skills: The Architecture Lesson I Got Wrong (and You'll Avoid)
Here's something the Hermes docs don't hammer hard enough: memory is an expensive tool and most people use it wrong.
The Mistake: Treating Memory Like a Catch-All Knowledge Base
I migrated from Claude Code to Hermes and immediately started dumping everything into memory:
- Provider setup procedures
- Stakeholder memo templates
- Plugin system architecture docs
- Markdown link formatting rules
- Token budget calculations
I ended up with ~6,000 characters of memory. Every response injected all 6KB into the system context. Result: token bleed.
Real cost: I burned money in a short session recalibrating a tracking plugin, largely because memory was bloated and the plugin was reading stale data that lived in memory.
The Architecture (Done Right)
Memory should ONLY contain persistent operating facts that change how I operate every single turn:
- Your timezone, written once and used everywhere
- Paths you refuse to let the agent guess
- A hard monthly token budget, checked before heavy sessions
- Recent failure patterns (plugin caches stale modules; file changes do not take effect mid-session; fully quit the app)
Everything else should be a skill:
- Provider setup procedures → skill (
migrate-to-hermes) - Stakeholder memo format → skill (
deliverables-format) - Plugin debugging checklist → skill (
hermes-plugins) - Markdown link rules → skill (
hermes-writing)
The Math
I had memory at 6,000 chars. Consolidated to 1,500 chars of actual operating facts. Moved the rest to skills.
Token cost reduction:
- Every response injects memory into context
- 6KB vs 1.5KB = 4.5KB saved per turn
- At ~1 char = 0.25 tokens, that's ~1,100 tokens per response
- If I do 10 responses a day = 11,000 tokens/day saved
- At Haiku pricing (~$0.80/M input), that's ~$0.009/day saved
- Over a month = ~$0.27 saved
That doesn't sound like much, but: add stale memory + bloated system prompts + repeated context injection, and token waste compounds fast. Skills don't inject unless loaded. Memory always injects. When you're on a tight budget, that matters.
What I Should Have Done Differently
Day 1 of migration: Ask the question upfront. Is this a fact that affects every single response, or is it a procedure I use sometimes?
- Fact → memory (at 500 chars per fact max)
- Procedure → skill
The cowardly move is to dump every procedure into memory so it is always there. The right move is: build a skill, test it, deploy it, and use memory only for the operating facts that cannot live in a skill.
The Implementation
Memory limit: 8,000 chars (plenty of room for 5-6 operating facts)
Current memory (lean):
- Provider auth rules (subscription is not API access)
- Token budget math
- Paths the agent must not guess
- Plugin failure pattern (stale module cache)
- One timezone rule
Everything else: skills.
FAQ
What is Hermes Agent?
Hermes Agent is an open-source AI agent framework by Nous Research. It runs in a terminal, a desktop app, and messaging platforms. It works with many model providers. It is not a ChatGPT wrapper. It is an agent runtime with skills, memory, hooks, and profiles.
How do I install Hermes Agent?
On macOS or Linux, run the official installer, then hermes setup or hermes setup --portal. The docs are at hermes-agent.nousresearch.com/docs. Windows has a PowerShell installer. After install, run hermes doctor if something fails.
Does a ChatGPT or Claude subscription work with Hermes?
Usually no. A ChatGPT Plus or Claude Pro web plan is not API access. Nous Portal is the exception: the subscription is the API. OpenAI Codex OAuth and Anthropic Max OAuth are documented paths for those products. OpenRouter is the easy path if you want many models with one key. Check the current quickstart. These options change.
What is the difference between Hermes Agent and Claude Code?
Claude Code is Anthropic's coding agent. Hermes is provider-agnostic. It adds persistent memory, skills, profiles, a messaging gateway, and a hook system that can rewrite output in code. You can import some Claude Code config with hermes import-agent. The hard part is mapping operating rules into memory and skills.
What is the difference between a Hermes nudge and enforcement?
A pre_llm_call hook injects a reminder before the model answers. The model can ignore it. A transform_llm_output hook runs after the model finishes and can rewrite the reply in code. The model cannot skip that. Nudge with the first. Enforce with the second.
Why did my Hermes plugin not change after I edited it?
The desktop app caches imported Python modules for the life of the process. A new session is not a restart. Quit the app fully, delete the plugin __pycache__, and relaunch.
Why does transform_llm_output not show in the app?
Turn streaming off: hermes config set display.streaming false. If streaming is on, the original text is already on screen before the hook runs. Also use a Python plugin, not a shell hook, unless your Hermes version has fixed issue #67890.
Should I put procedures in Hermes memory or in skills?
Memory is injected every turn. Keep it to operating facts: timezone, budget, paths you refuse to guess, recent footguns. Put procedures in skills. Skills load when they are relevant. Memory always costs tokens.
How do I run Hermes on a free model?
Point Hermes at OpenRouter and pick a live :free model, or run a local model through Ollama. Free cloud tiers often allow training on your prompts. Do not send private code through those tiers. Local weights keep the data on your machine. See Free AI Models for Hermes Agent.
Sources
- Nous Research. Hermes Agent documentation. https://hermes-agent.nousresearch.com/docs/. The docs are the source of truth for commands and APIs. This field guide is the operational layer the docs do not assemble for you.
- Nous Research. Hermes Agent GitHub. https://github.com/nousresearch/hermes-agent.
- Nous Research. Hermes Agent quickstart. https://hermes-agent.nousresearch.com/docs/getting-started/quickstart.
- OpenRouter. "Hermes Agent + OpenRouter." https://openrouter.ai/blog/tutorials/hermes-agent/. Setup and routing. I do not copy that tutorial.
- OpenRouter. Model listings and free-tier terms. https://openrouter.ai. Verify live prices and
:freestatus; they change. - Hermes Agent issue #67890. https://github.com/NousResearch/hermes-agent/issues/67890. Shell-hook
transform_llm_outputdropped stdout as of July 2026. Re-check before you depend on the shell path. - Lessons in this guide come from a personal migration off Claude Code onto Hermes. I ran it on personal accounts, in off hours.