EvoMap
Hermes Agent Security Guide for Self-Hosting

Hermes Agent Security Guide for Self-Hosting

April 29, 2026
1,040 views
hermes-agent security self-hosting docker mcp agent-ops

Hey, I'm Lena. I spent the better part of an afternoon reading the security pages before I touched a single config file. Then I went back and read them again. The first time through, the system felt like a long list of options. The second time, I started to see the shape of the thing — there isn't really a "security feature" in ​Hermes​ Agent security​, there's a stack of independent layers that each assume the others might fail.

That framing matters more than I thought it would. Most agent security writeups I've come across read like a checklist of tips — turn this on, set that flag, you're fine. The way Hermes is actually documented is different. It's a defense-in-depth model where every layer is doing one specific job, and turning any single one on doesn't substitute for the others. No layer is the whole answer. That sentence is basically the whole post.

This is a write-up of how I'm thinking about hardening a long-running self-hosted Hermes deployment. I'm still figuring some of it out.

The Hermes security model in layers

If you read the official security page carefully, the model has roughly seven independent layers — and the word "independent" is doing real work there. Each layer makes a different assumption about what could go wrong. Listed in roughly the order they apply during a tool call:

  • User authorization — gateway-level check on who can talk to the agent at all
  • Dangerous command approval — pattern-based detection on shell commands before they run
  • Tirith content scanner — Rust-based pre-execution scan for prompt-injection, credential exfil, terminal-injection patterns
  • Runtime isolation — local vs Docker vs SSH vs Modal vs Daytona vs Singularity
  • Credential filtering — env-var stripping for child processes, MCP subprocess env allowlists, output redaction
  • Context file scanning — AGENTS.md, SOUL.md, .cursorrules scanned for prompt injection at load time
  • Cross-session isolation — sessions can't read each other's state; cron paths hardened against traversal

I keep coming back to this because the layer that catches a given problem is not always the layer you think it should be. A malicious shell command might be caught by approval, by Tirith, or by the container — or all three, or sometimes none. That's why turning one off feels safe right up until it isn't.

The principle this is borrowing from is older than agents. CISA's Secure by Design joint guidance talks about secure defaults — the idea that the out-of-the-box configuration should be the safe one, with deviation being explicit. Hermes mostly follows this: approvals are on by default, Tirith fails closed in high-security mode, gateway denies all users when no allowlist is set. The catch is that ​mostly​, and the deviations matter.

Dangerous command approval and where it helps

Here's the part I want to be careful about. Approval prompts feel reassuring, and that's exactly why they're easy to over-trust.

Hermes maintains a list of regex patterns for dangerous commands — things like rm -rf, DROP TABLE, fork bombs, piping curl output to bash, killing the gateway process. When a match is hit, the user is asked to approve. Three modes: manual (always ask), smart (LLM-assisted risk scoring), off (don't ask). There's also /yolo for the session that bypasses everything.

The thing I keep reminding myself: regex-based detection is fundamentally bypassable. A command can be obfuscated, base64-encoded, split across two tool calls, written to a file and then sourced. Hermes does normalize input (strips ANSI escapes, null bytes, NFKC Unicode) before scanning, which closes the obvious obfuscation paths. But it's still pattern matching against an open-ended attack surface.

A community security audit on the Hermes GitHub recently went through this in detail. The finding wasn't that Hermes is unsafe — it was that the default configuration is permissive and assumes the user will harden it. That's a defensible design choice. It's also one that puts real responsibility on whoever is running the deployment.

So the way I think about approval now: it's a tripwire for accidents, not a defense against an adversarial agent. When a model with tool access goes off-script in a way that's unintentional — wrong path, missing flag, a confidently wrong command — approval catches it. When something has actually compromised the model's input, approval is the wrong layer to be relying on. The approval prompt is the ​human-in-the-loop​​ seatbelt; the container is the crumple zone.

Runtime isolation choices

This is where Hermes gives you real options, and the choice has more weight than the others. The terminal backend determines what "the agent ran a command" actually means in physical terms.

  • local — runs commands on the host. Default. Approval prompts active. The blast ​radius​​ is the host.
  • docker — commands run inside a container with read-only root, dropped capabilities, configurable CPU/memory/disk. Dangerous command checks are unconditionally skipped because the container itself is the boundary.
  • ssh — runs commands on a remote machine. Useful for keeping the agent off your laptop entirely.
  • modal and daytona — serverless backends that hibernate when idle. Container-class isolation, near-zero idle cost.
  • singularity — for HPC environments where Docker isn't available.

A couple of things worth pulling out. The "container bypass" of approval prompts is intentional — the team's reasoning is that if the container holds, the inner regex check is redundant; if the container fails, the regex wasn't going to save you anyway. I had to sit with that for a minute before I agreed with it.

If you do go Docker, the Hermes config exposes hardening levers that map closely to the standard practices in Docker's official engine security docs — namespace isolation, dropped capabilities, read-only root, resource limits. The default Hermes Docker image runs with these on. The mistake people make is adding things back. terminal.docker_forward_env is the obvious one: every variable you forward into the container is one the agent can read and exfiltrate. Empty allowlist by default, and I'd keep it that way for anything but task-specific tokens.

Persistent vs ephemeral mode is another fork. Persistent bind-mounts a workspace directory across runs, so the agent can build state — useful for development. Ephemeral uses tmpfs, so everything dies when the container stops — useful for anything closer to production. Pick based on whether you'd be okay with a malicious file living in ~/.hermes/sandboxes/ for a week before you noticed.

Messaging and MCP security boundaries

When Hermes runs as a gateway, "who can talk to it" becomes a real question. The authorization model is layered: per-platform allow-all flags, DM pairing approved list, platform allowlists, global allowlist. The default if none of these are configured is ​deny everyone​, with a startup warning. That's the right default — and the fact that some self-host guides walk users through setting GATEWAY_ALLOW_ALL_USERS=true for testing and then forgetting to undo it is its own category of incident.

DM pairing is the part I keep recommending to people. Instead of you maintaining a list of Telegram/Discord IDs, an unknown user gets a one-time pairing code, and you approve it from the CLI with hermes pairing approve telegram ABC12DEF. The Hermes docs note this design draws on OWASP and NIST SP 800-63 digital identity guidance — codes are one-time, time-bounded, and the approval action is explicit. It's not novel, but it does mean access decisions are made with a real human looking at the request, not by editing config files in advance.

MCP is its own surface entirely. The thing that took me longest to internalize: an MCP tool inside ​Hermes​​ runs with whatever permissions the server it lives behind has, not the agent's. That's the whole point of the protocol, but it also means every connected MCP server is a separate trust decision.

Hermes' MCP config reference documents two protections that I now consider non-optional. First, environment filtering: by default only PATH, HOME, USER, LANG, LC_ALL, TERM, SHELL, TMPDIR, and XDG_* are passed through to MCP stdio subprocesses. Everything else — API keys, tokens — is stripped. Variables you actually need go in the server's explicit env block. Second, per-server tool filtering with include and exclude: if a server exposes 20 tools and the agent needs 3, allowlist the 3.

The third one isn't config — it's discipline. Context files (AGENTS.md, SOUL.md, .cursorrules) are scanned for prompt injection patterns before being loaded into the system prompt. This is the layer that addresses what OWASP's LLM Top 10 calls indirect prompt injection — instructions hidden inside content the model is told to read. The scanner isn't a complete defense (nothing is, in this category), but it's the right place for that check to live, and it's on by default.

Hardening a self-hosted Hermes deployment

A few patterns I've ended up using, in rough priority order:

Run as a non-root user with passwordless ​sudo​​ only where required. The Hermes installer assumes this; the security model assumes this. Don't run the gateway as root.

Pick the runtime isolation that matches your trust model, not your convenience. For a server that's mostly running scheduled jobs and answering messages, Docker backend with empty docker_forward_env, ephemeral workspace, and minimal mounted volumes is the boring correct choice. Local backend is reasonable on a personal laptop where you're watching every approval prompt; it's a worse default on a VPS you're not actively staring at.

Keep secrets out of broad scopes. Provider API keys and gateway tokens go in ~/.hermes/.env with 0600 permissions. They're never in env_passthrough. They're never in docker_forward_env. MCP servers get only the variables their env block explicitly names.

Configure user authorization explicitly. Either set platform allowlists with the user IDs you actually want, or use DM pairing. Never leave ​GATEWAY_ALLOW_ALL_USERS=true​​ in a production environment — it's a testing flag.

Audit the layers you've turned off. approvals.mode: off and /yolo exist for a reason — CI runs, sandboxed test environments — but if either is set in your live config, write down why. Same for tirith_enabled: false. The value isn't whether they're on; it's whether you can answer why they aren't.

Update discipline matters more than the initial config. Hermes ships frequently, and security improvements have landed in nearly every recent release — secret-exfil blocking, expanded credential directory protections, broader token redaction patterns, browser URL exfil blocks. A six-month-old install is a worse install regardless of how carefully you set it up.

I'll keep refining this. Some of the layers I still don't fully understand — Tirith's verdict-to-approval handoff is one I want to spend more time with, and I haven't done a real audit pass on a long-running deployment of my own. What I'm reasonably sure of is that the layered model is right, and that no individual layer is the whole answer. If a writeup tries to hand you one trick, it's probably wrong.

FAQ

​Should I run with ​approvals.mode: off​?

Only inside a container or sandbox where the container is doing the work. On the local backend, no.

Is the Docker backend a complete substitute for approval prompts?

For host-level damage, mostly yes — that's why approval is bypassed when a container backend is active. It is not a substitute for credential filtering or MCP boundaries.

What does ​YOLO​​ mode actually disable?

Dangerous command approval prompts for the current session. It does not disable Tirith, container isolation, gateway authorization, or env filtering.

How do I know which ​env​​ vars ​MCP​​ subprocesses see?

Only PATH, HOME, USER, LANG, LC_ALL, TERM, SHELL, TMPDIR, XDG_*, and anything in the server's explicit env block. Everything else is stripped.

Is there a managed ​Hermes​​ that handles this for me?

There are providers offering one-click deployments. They handle the install — they don't make the trust decisions about who can talk to your bot or what your MCP servers can read. Those still have to be yours.

I'll come back to this when I've spent more time with my own setup. For now, this is where my understanding is.

Previous Posts:

Related Articles