The problem: agents forget everything between sessions
A language model has no memory. Every conversation starts from zero; everything it "knows" about you must fit inside the context window of the current session. For a chatbot that answers one question, that is fine. For an agent that works for you daily — managing bookings, following up leads, tracking projects — amnesia is fatal: it re-asks what it already learned, repeats mistakes, and cannot build on yesterday’s work.
That is why agent memory management has become one of the most-watched topics in AI engineering, and why searches for Anthropic’s approach spiked after the company published how its production agents handle it.
It is worth separating two things that get conflated. A long context window is not memory — it is a bigger desk. Everything on the desk is cleared when the session ends, and a bigger desk costs more to maintain every single turn. Memory is what survives the desk being cleared, and the engineering question is what deserves to survive.
Context window versus memory — the distinction that matters
Every turn of a conversation re-sends the entire working context to the model. That has two consequences people underestimate.
The first is cost. If you keep a 200,000-token history in context, you pay to re-transmit it on every message, not once. A long-running agent that hoards its history has an operating cost that grows quadratically with the length of the relationship, which is exactly backwards from what you want.
The second is quality. Models attend less reliably to material buried in a very long context than to the same material presented compactly. An agent that carries three months of raw transcript is not better informed than one carrying a curated page of durable facts — it is usually worse, because the signal is diluted.
This is why the serious designs treat context as a scarce resource to be curated rather than a bucket to be filled, and push durable knowledge out into memory that is loaded selectively.
How Anthropic’s approach works, in plain English
Anthropic’s design, described across its engineering posts and agent documentation, rests on three ideas. First, memory as files the agent maintains itself: instead of a mysterious vector database, the agent writes and edits plain files — an index plus one file per topic — recording durable facts, preferences, and project state. Before each task it recalls what is relevant; after meaningful events it updates the files. Memory becomes inspectable: you can open it, read it, and correct it.
Second, context engineering over context hoarding: because the context window is finite and expensive, the agent curates what enters it — summarizing finished work, discarding stale tool output, keeping durable knowledge in memory files rather than re-reading full transcripts. Anthropic calls the discipline "effective context engineering"; compaction and context editing are the mechanisms.
Third, consolidation — the part headlines nicknamed "dreaming": between tasks, the agent reviews accumulated raw memory and rewrites it — merging duplicates, promoting patterns into stable knowledge, pruning what no longer matters. Like human sleep, it turns a messy log of experiences into usable long-term memory, which is what lets long-running agents improve instead of drowning in their own history.
Why files instead of a vector database
The instinct of most teams building agent memory is to reach for embeddings and a vector store. It is worth understanding why a file-based approach earns its place, because the trade-off is not obvious.
Vector search answers "what in my history is semantically similar to this question?" That is powerful for retrieval over large document collections, and genuinely the right tool there. But agent memory is a different shape of problem: it is usually a small body of facts that must be exactly right, and where being approximately right is worse than useless. "The client invoices on the 5th" is not something you want retrieved by similarity — you want it read.
Files also give you three properties that matter operationally. They are inspectable: a human can open the memory and read what the agent believes. They are correctable: if a fact is wrong, you edit the line. And they are portable: the memory is not locked inside an index whose format belongs to a vendor. When something goes wrong with an agent — and something will — being able to read its beliefs is the difference between a fix and a shrug.
The honest limitation: files do not scale to hundreds of megabytes of knowledge. When memory outgrows what can be indexed and selectively loaded, retrieval becomes necessary. Most business agents never reach that point.
Memory as current state, not as a log
This is the single rule that separates memory that helps from memory that hurts, and it is the one most implementations get wrong.
The natural instinct is to append. Each session, the agent adds what happened. The file grows. Nothing is ever removed because removing feels like forgetting. Within weeks the memory is a chronological log, and every task begins by reading it in full.
We measured exactly this failure on a production deployment: a single agent’s memory directory had reached 492 KB — roughly 115,000 tokens — re-read on every scheduled run, with the largest single file at 165 KB. The agents were treating memory as a journal, writing one new entry per run and pruning nothing. It was the largest single line item in that account’s token cost, and none of it was making the agent smarter.
The correct model is that memory holds current state. "The client prefers morning appointments" replaces the previous belief rather than joining it. A project entry reflects where the project is now, not every step it took to get there. History belongs in a transcript that can be searched on demand; memory holds what the agent should know without looking anything up.
What belongs in memory — and what must never
A workable rule: memory holds facts that are durable, specific and would change the agent’s behaviour if it knew them.
Durable means it will still be true next month. Someone’s role, a recurring schedule, a system the business uses, a preference expressed more than once, a constraint that keeps applying. Not the content of Tuesday’s conversation.
Specific means it is actionable. "The client is busy" helps nobody. "The client does not answer calls before 10am and prefers WhatsApp" changes what the agent does tomorrow.
Behaviour-changing is the sharpest filter. If knowing a fact would not alter a single decision, it is trivia, and trivia costs the same to carry as knowledge.
And the hard prohibition: no credentials, no card numbers, no documents, no passwords — ever, under any circumstance, regardless of how convenient it would be. Memory is read into context on every task and is far more exposed than a credential store. When we audited the memory directory mentioned above, we found a door code, a room number, a phone number and a spreadsheet identifier sitting in plain text — all of it written by agents whose own instructions prohibited exactly that. The lesson is that a policy in a prompt is not a control. Secrets need to be blocked by code that refuses to write them, not by an instruction asking nicely.
Isolation: whose memory is it?
The moment more than one client or more than one agent shares a system, memory isolation stops being hygiene and becomes the whole ballgame.
Two boundaries need to hold. Between customers, obviously — one tenant’s agent must be structurally incapable of reading another’s memory, enforced by the filesystem and access controls rather than by the agent choosing not to look. And between agents belonging to the same owner, which is the one people skip: if every agent in an account shares one memory pool, a monitoring agent inherits the state of a posting agent, "clear this agent’s memory" becomes impossible without wiping the others, and each agent pays context cost for knowledge belonging to jobs it does not do.
We learned the second one the expensive way. When we moved to per-agent memory, we initially seeded each new agent from the shared pool so nothing would be lost in the transition. It was the wrong call: agents woke up holding another agent’s working state — a posting agent carrying a monitor’s notes about an unrelated system. We removed the seeding. A new agent now starts empty, and adopting anything from a legacy pool is a deliberate, file-by-file choice made by the owner. Convenience is not worth contamination.
Keeping memory from growing without bound
Because memory enters context on every run, its size is a recurring cost, which means it needs a ceiling and an enforcement mechanism rather than a guideline.
A two-tier limit works well in practice. A soft limit, where the agent is told its memory is getting large and instructed to consolidate before doing anything else — merge duplicates, replace superseded facts, delete what stopped mattering. And a hard limit, where the system archives the offending file automatically and leaves a stub pointing at the archive. Archive rather than delete: recoverable, but out of the context path.
The consolidation instruction should be explicit about the model: memory is current state, substitute rather than append, never store secrets. Left to their own devices, agents append — it is the lower-effort action and nothing in their training pushes back on it.
Layering memory: session, rolling state, transcript, permanent facts
A single memory store is rarely the right architecture, because different knowledge has different lifetimes. Four layers cover the realistic cases.
The live session is the current conversation, held in context, which ends when the session does. It is fast and free-form and it should not be relied on for anything, because sessions are cut by timeouts, size limits and restarts far more often than users expect.
A rolling state is a short summary of where this particular conversation stands, rewritten after each exchange and injected on every turn. It is what makes a conversation survive its own session ending — the agent picks up mid-thread even though the underlying session is new.
A searchable transcript is the full history, retained for a defined period, queried on demand rather than loaded by default. This is where "what did we decide last Tuesday" gets answered, without paying for Tuesday on every unrelated turn.
Permanent facts are the small, curated set that is true independent of any conversation, loaded always. This is memory proper, and it should stay small enough to read in a minute.
The reason to separate them is that each layer has a different retention policy, a different cost profile, and a different failure mode. Collapsing them into one store means the shortest-lived data drives the cost of the longest-lived.
Why memory management decides whether your AI agent is useful
Memory is not a research curiosity — it is the difference between a demo and an employee. An agent with managed memory remembers that invoices go out on the 5th, that a given client always reschedules, which supplier answered last time.
The compounding is the point. On day 1 an agent knows what you told it. On day 90, if memory is working, it knows your business — the exceptions, the preferences, the recurring patterns that nobody ever wrote down. That accumulated context is what makes the difference between an assistant you have to brief every time and one you can simply ask.
And it only compounds if the memory stays clean. An agent that has accumulated 90 days of undifferentiated log is not more knowledgeable than on day 1; it is slower, more expensive, and more likely to act on something that stopped being true in week three.
A checklist for evaluating agent memory
Whether you are building or buying, seven questions separate real memory management from a marketing bullet.
Can I see the memory? If you cannot read what the agent believes, you cannot correct it. Can I edit or delete a specific fact, without wiping everything? Is memory isolated per customer and per agent, enforced by the system rather than by convention? Is there a size limit, and what happens when it is hit? Does the agent consolidate, or only append? Are secrets blocked by code rather than by instruction? And how long is the underlying transcript kept, separately from the memory itself?
We apply these exact principles at Genesis AI Labs: every AI employee we deploy keeps per-agent persistent memory — isolated per client, size-capped with automatic archiving, inspectable and cleanable from a dashboard, layered so that a cut session never costs the thread, with secrets banned by policy and blocked in code. Your agent on day 90 knows your business far better than on day 1, and you can audit every fact it holds. That compounding is where the ROI of AI agents actually lives.