Hi, I'm Lena. I had a question I kept putting off about Hermes Agent memory. Every time someone described it as "the agent that learns," I paused for a moment. Learns what, exactly? And remember where it was? I went back to the docs a few times before I felt I had even half an answer. This is what I noticed.
I'm writing this not as someone who built Hermes, but as someone who watched the system long enough to see where its memory does what people expect — and where it quietly doesn't. The gap between those two is bigger than the marketing language suggests, and it's worth being honest about.
What Hermes memory actually stores
The first thing I had to unlearn: Hermes doesn't have one memory system. It has several, and they don't do the same job. Most confusion I've seen comes from treating them as one thing.
MEMORY.md, USER.md, and frozen session snapshots
The built-in layer is two markdown files living in ~/.hermes/memories/. According to the official Hermes memory documentation, MEMORY.md is capped at around 2,200 characters (roughly 800 tokens) and holds environment facts, project conventions, and lessons the agent thinks are worth keeping. USER.md is smaller — about 1,375 characters, around 500 tokens — and holds preferences about you.
Both are loaded at session start as a frozen snapshot injected into the system prompt. That word — frozen — is the part I kept missing on the first read. Mid-session writes are persisted to disk immediately, but the snapshot in the active prompt doesn't refresh until the next session begins.
I sat with this for a moment. It explains a behaviour I'd seen and couldn't place: the agent saves something, says it saved it, and then proceeds as if it doesn't quite have it yet. That isn't a bug — it's the snapshot model. Once I understood that, a few small inconsistencies stopped feeling inconsistent.
What "persistent" means in practice
"Persistent" here is doing real work, but a narrower kind of work than the marketing word suggests. The files survive between sessions. They are agent-curated, not agent-recorded — the LLM decides what's worth saving. And they're capped on purpose: a small, stable system prompt is what makes prefix caching efficient, and prefix caching is part of why the agent feels responsive.
So when someone says Hermes "remembers," what they usually mean, more precisely: a small, deliberately bounded notebook is reloaded at the top of each session, and a separate searchable archive of past conversations exists alongside it. Two different things, with two different access patterns.
What Hermes memory is good at
I want to be fair here. The built-in system is genuinely useful — it just isn't what most people initially imagine.
Preferences, environment facts, recurring context
The sweet spot is small, durable, frequently-relevant facts. Things like:
- "User's project is a Rust web service using Axum + SQLx"
- "User prefers concise responses, dislikes verbose explanations"
- "This machine runs Ubuntu 22.04, has Docker installed"
That kind of context belongs in the system prompt every time, and Hermes handles it well. The agent saves automatically when it judges something worth keeping, and it consolidates when entries pile up — merging three "project uses X" lines into a single project description, for example. In short or narrowly-focused sessions you may see nothing written at all — that's a feature, not a failure. Most short tasks shouldn't be polluting your long-term notebook.
There's also a security-scanning step on memory entries to catch prompt-injection attempts, which I appreciated noticing in the Hermes GitHub source. It's a small detail, but the kind of detail that tells you someone thought about what could go wrong when an LLM is in charge of writing its own memory.
What Hermes memory does not solve
This is where I had to slow down. A lot of expectations break here, and I don't think it's because the system is doing something wrong — I think it's because the word memory is doing too much work.
Bounded memory, session resets, no automatic validation of learned behavior
A few things are worth being honest about:
Memory is bounded. ~2,200 chars for environment, ~1,375 for user. When the limit is reached, the agent has to consolidate or remove entries before adding new ones. Nuanced specifics can get compressed away in that process. This is the single most common surprise — people assume memory grows; it doesn't. A small, fixed budget is the design.
Mid-session writes don't appear in the same session. The frozen snapshot model means the agent acts on what it loaded at start, even after writing new entries. Restart the session and they're in context. If you need the agent to act on something it just wrote, reference it explicitly in the conversation — don't expect it to appear automatically.
The agent has to decide to save. Built-in memory is judgment-based. There's no automatic transcription. A configurable nudge_interval periodically prompts the agent to reflect, but in short sessions you may genuinely see empty files. I might be reading too much into this, but I think it's the most under-communicated part of the design.
There is no automatic validation. If the agent saves a "lesson learned" and the lesson was actually wrong, nothing flags that. The fact stays in the snapshot until something replaces it. Memory is not the same as verified knowledge — it's just what one model thought, at one point, was worth keeping.
Memory is not the same as reusable capability
This is the part I had to reread my own notes on. Hermes also has skills — markdown documents in ~/.hermes/skills/ that capture procedures, tools used, and steps that worked. Skills are created reactively after complex tasks (typically 5+ tool calls) and load on-demand using progressive disclosure: Level 0 is just a list of skill names and short descriptions, Level 1 loads the full content of a specific skill when needed.
Memory and skills are different things, even though both look like markdown files.
Memory is "what this user prefers and what environment we're in." Skills are "how to do this kind of task." A preference doesn't tell the agent how to debug an OAuth flow; a skill might. Conflating the two is, I think, the single biggest source of confusion when people say "the agent isn't learning." It might be learning — just not in the layer they're checking.
Cross-session recall, session search, and where people get confused
Beyond MEMORY.md and USER.md, Hermes stores all CLI and messaging sessions in SQLite at ~/.hermes/state.db with FTS5 full-text search. The agent can call a session_search tool to retrieve past conversations, which are then summarized using a small model.
Two details worth knowing.
FTS5 is keyword-based. It's a powerful and well-engineered index — the SQLite FTS5 documentation describes how it tokenizes content and matches on those tokens — but it matches on exact tokens, not meaning. If a past session said "authentication microservice uses Redis," asking "what did I tell you about the auth service?" may not retrieve it. The agent has to know to call session_search, and use the right query terms. There's no entity resolution, no relationship tracking, no semantic rephrasing.
Session search isn't automatic. It's a tool the agent decides to call. If the agent doesn't think to call it before answering, the past conversation isn't consulted. That's a behaviour gap, not a storage gap — the data is there, the retrieval just didn't happen.
This is where I see people get most frustrated. They saw the agent discuss something three weeks ago and assume it'll come up naturally. Sometimes it does. Often it doesn't, because nothing triggered the search. Vectorize has a useful breakdown of these failure modes in their troubleshooting article on Hermes memory that I went back to twice.
Practical design advice for long-running Hermes workflows
I'm hesitant to give advice — I'm one data point. But a few things stood out from watching this for a while.
Treat built-in memory as a small notebook, not a database. Anything that needs to scale with conversation length doesn't belong there. Use it for stable facts that should always be in context, and accept that detail will get compressed when the budget runs out.
Be explicit when something matters. "Remember that my production database runs on port 5433" works better than hoping the agent flags it on its own. Built-in memory is curated, not recorded — and the agent's judgment about what matters won't always match yours.
Reach for an external provider when you need structured recall. The Hermes memory providers page lists eight pluggable options — Honcho, OpenViking, Mem0, Hindsight, Holographic, RetainDB, ByteRover, and Supermemory. They sit additively on top of MEMORY.md and USER.md, which keep working unchanged. The built-in layer doesn't go away; the external one fills in what it can't do.
For cross-session user modelling specifically, Honcho takes a peer-based approach — user and AI are both modelled as peers, each with its own representation that updates from observations over time. It's a different design philosophy from a flat fact store, and I'm still working out where I'd reach for it versus a simpler vector store. But the multi-pass dialectic reasoning is interesting, especially the cold-start versus warm-session distinction.
Don't confuse memory with capability. If you want the agent to do something better next time, you probably want a skill, not a memory entry. And if you want it to be correct next time, neither memory nor skills will validate that for you — you have to. That's the part I keep coming back to.
FAQ
Does Hermes Agent memory get smarter over time?
The files accumulate facts. Whether that constitutes "smarter" depends on what you mean. The agent uses what's in the snapshot, but it doesn't reason over the history of memory edits unless you've added an external provider that does. I'd be careful with the word learn here.
Why is MEMORY.md empty after several sessions?
Most likely the agent never judged anything worth saving — short or task-focused sessions often produce no writes. Check your nudge_interval config, or reach for an external provider if you want automatic capture without relying on the agent's judgment.
Can I just make the limit bigger?
You can. The character limits are configurable in ~/.hermes/config.yaml. But the cap exists for a reason — bigger snapshot, less room for actual conversation, and weaker prefix-cache benefits. Bigger isn't free.
What happens to memory mid-session?
Writes go to disk. The active prompt doesn't refresh until the next session. This is expected, and trying to rely on within-session memory updates will cost you debugging time.
Is session search the same as memory?
No. Memory is the always-loaded snapshot. Session search is an on-demand keyword lookup over past conversations. Different mechanisms, different latency, different failure modes.
I'm not ready to close this one yet. There's a lot in the Hermes Agent memory stack I want to keep watching — especially how external providers behave once a profile has months of data behind it. For now, this is where my understanding sits.
Previous Posts:
- 👉 Understand how Hermes memory compares to broader agent memory architectures
- 👉 Explore why agents struggle with consistent learning across sessions
- 👉 Learn the real difference between agent skills and reusable capabilities
- 👉 See how self-evolving skills improve long-term agent performance
- 👉 Understand how agent assets differ from simple memory storage




