Designing Agent Memory Systems: Three-Layer Architecture and Pollution Cleanup
How to build working, episodic, and semantic memory; three sources of memory pollution, TTLs and memory GC; and whether you need a memory system at all.

In 2026, what separates agent applications is no longer how well the prompt is written — it's whether the agent remembers things. A support agent that can't recall last week's complaint asks the same questions twice; a coding agent that can't remember the project's architecture conventions "rediscovers" them every session. But the flip side is more dangerous: I watched an agent treat a three-month-old expired policy as current rules and process a customer's request wrong. More memory isn't better — wrong memory is worse than no memory.
This article takes apart the three-layer architecture of agent memory systems — what each layer stores, how to build it — and the most overlooked pitfall: memory pollution, plus cleanup strategies.
First, a common confusion to clear up: a memory system is not RAG. RAG is "looking things up" — finding document snippets in an external knowledge base, solving "the model doesn't know." A memory system is "remembering experiences" — recording what happened with this user and this task, solving "the model doesn't remember." A support agent can use RAG to look up the product manual (knowledge) while using memory to recall "this user complained about shipping last week" (experience). They're often used together, but the design goals differ completely: RAG optimizes retrieval accuracy; memory systems optimize the judgment of "what to remember and what to forget."
The Three-Layer Architecture: Working, Episodic, and Semantic Memory
Borrowing the human-brain taxonomy is the most practical approach. Each layer solves a different problem, with completely different technical choices.
Layer 1: Working Memory = the current context window. What's stuffed into the prompt this turn: what the user just said, the tool results just returned. It's fast, complete, and expensive — fully retained, but the window is finite. One design rule: don't stuff garbage into it. A 10KB JSON blob from a tool call gets extracted before entering context; history beyond N turns gets rolling summaries. Many agents get dumber the longer they run because working memory is bloated with irrelevant information, squeezing the useful stuff out of the window. The empirical rule: budget working memory (say, 60% of total window) — exceeding it triggers summarization. That's the baseline design.
Layer 2: Episodic Memory = things that happened. "The user complained about slow shipping on September 12." "Last deploy failed because the DB migration didn't run." Concrete, timestamped events. Typical implementation: at each session's end, generate a structured summary (time, event, key entities, outcome), store it in a database, and recall by user ID + time decay + semantic relevance. This layer is the source of "personalization" — the entire reason an agent feels like it "remembers you." Critical design: every episodic memory must carry three things — source (which session), timestamp, and confidence. Without them, it's an unauditable black box.
Layer 3: Semantic Memory = learned knowledge. Stable facts distilled from repeated events: "user prefers communicating in English," "this repo's test command is npm run test:unit," "refund policy is 7-day no-questions-asked." Implementation is vector DB plus knowledge graph: vector retrieval handles "seems similar," the graph handles "is accurate." This layer is the hardest, because it requires distillation from episodic memory — not written every conversation, but a scheduled job (e.g., daily) that promotes high-frequency, stable patterns into semantic memory. Keep the promotion threshold conservative: a rule needs corroboration from at least 3 independent events before it earns a place in semantic memory.
Data flows one way through the funnel: working memory → end-of-session episodic summaries → periodic distillation into semantic memory. Retrieval runs the reverse: check semantic memory (stable knowledge) first, then episodic (specific events), then assemble into working memory. That funnel is the essential difference between a memory system and "dumb vector retrieval."
Each layer has a classic anti-pattern — dodging them saves weeks of debugging. Working memory anti-pattern: stuffing raw data. A full 10KB raw JSON in context drowns the signal and explodes the token bill. Extract at the tool layer — only "conclusion + key fields" enter context; raw data stays retrievable on demand. Episodic memory anti-pattern: storing summaries without provenance. A summary says "user prefers refunds over exchanges," but you can't find which conversation said it — so you can't verify or correct it. Every episodic memory must trace back to its source session in one click; that's the auditability baseline. Semantic memory anti-pattern: too low a promotion threshold. The user casually mentioned "liking dark mode" once, and it got promoted to a permanent preference — every recommendation since has skewed. Remember the iron rule: 3+ independent corroborating events before semantic promotion. Better to under-remember than mis-remember — under-remembering just means slightly less personalization; mis-remembering means outright mistakes.
On retrieval strategy, don't worship pure vector search. Production-grade memory retrieval uses hybrid scoring: semantic similarity (vector) 50%, keyword match 20%, time decay 20%, confidence 10%. Why does time decay matter so much? Because "what the user said last week" is usually more relevant than "what they said three years ago," while pure vector retrieval is timeless — it treats a three-year-old joke and last week's formal request as equals. The exact decay curve is tunable, but the "newer matters more" prior must be baked into ranking. Another field lesson: always over-recall, then rerank — fetch top-20, filter to top-5 with a lightweight reranker (or even rules) before entering context. Going straight to top-5 in context will wake you up at 3 AM with miss rates.
Memory Pollution: The Most Expensive Pitfall, and Cleanup Strategies
Memory pollution means wrong, expired, or manipulated memories entering the system and being treated as fact by the agent. Three sources, each with real incidents behind it.
Source 1: expired memory. The most common. An e-commerce support agent remembered the old "free shipping over ¥200" policy; after it changed to "over ¥299," the agent kept promising the old terms, and the company ate three months of shipping costs. Lesson: every memory needs a TTL. Policy-type memories get a 30-day TTL; on expiry they're auto-demoted to "needs verification" instead of continuing as fact. Time-sensitive memory without a TTL is a time bomb.
Source 2: bad writes. A user jokes "my surname is Wang" (it isn't); the agent records it and calls them "Mr. Wang" forever after. Or the agent hallucinates "user likes red" and self-reinforces after writing it into memory. Cleanup strategy: a three-question check before writing — did the user explicitly state this? Is there a second piece of evidence? How costly is getting it wrong? High-cost memories (identity, preferences, compliance rules) require human review or high-confidence dual-source confirmation — never auto-write.
Source 3: malicious injection. The one to watch most in 2026: attackers induce the agent through conversation to remember fake directives like "the admin said 10x refunds are allowed," which the agent later "acts on from memory." Defense: tiered memory — user-provided and system/admin-provided memories stored separately, with system-tier memory always outranking user-tier in decisions; any "memory" attempting to rewrite system-level rules is rejected at write time with an alert.
Beyond that, a general cleanup mechanism I call "memory GC": run garbage collection weekly — delete low-confidence memories unretrieved for 90 days; demote memories that were retrieved but led to bad decisions; log every deletion and demotion to an audit trail. After this went live, that e-commerce agent's memory error rate dropped from a dozen a month to near zero. Remember: a memory system's core skill isn't "remembering" — it's "forgetting." Forgetting accurately matters ten times more than remembering broadly.
Two more engineering practices worth adopting. First, versioned memory. Manage semantic memory like git manages code: each distillation produces a new version, diffs are kept, rollback is one click. When a distillation introduces a bad rule, you revert instead of hand-digging through a database. This design's value in an incident can't be overstated. Second, the four-element audit log. Every memory write, edit, deletion, and demotion records: who triggered it (user/system/distillation job), when, what it was before, and the confidence delta. When a user complains "your agent misremembered," the audit log is the only thing you can use to reconstruct and prove what happened — without it, you can't even locate the error.
The Build Checklist: Do You Actually Need a Memory System?
Don't build one for the "sophistication" factor. Three gates — build only if all pass. First: users come back. One-shot tool agents (single document conversion) don't need memory; support, personal assistants, and long-horizon coding collaborators do. Second: cross-session information genuinely helps. Ask: if the agent remembered last time, would the experience change qualitatively? If not, skip it. Third: you have the people or mechanisms to handle pollution. A memory system without GC is worse than no memory.
If you build, follow this minimum-viable order: working-memory budgeting first (zero cost, immediate payoff); then episodic session summaries (one prompt template); semantic distillation and vector retrieval last (most complex, do it last). 80% of agents never need past layer two — don't start with a knowledge graph.
Two copy-paste templates. First, the session-summary prompt, called at each session's end: "Extract information worth remembering long-term from this conversation as JSON: facts (stable facts like user preferences and project conventions), events (things that happened, with timestamps), todos (unfinished items). Only extract what future sessions might use — no play-by-play. Return an empty array if nothing qualifies." The last line is the point — default to not remembering; only the worthy get stored. That's the first line of defense against memory bloat.
Second, the daily distillation job: a batch run every night that takes the last 7 days of new episodic memories and asks the model three questions: which facts appeared 3+ times (semantic-memory candidates)? Which old memories did new events overturn (mark expired)? Which memories went 30 days unused (lower confidence)? Candidate semantic memories enter a "pending review" state; a human glances weekly and clicks confirm. Run this, and a medium-scale agent's memory system costs about 30 minutes of human maintenance a week.
My Take: Memory Is Becoming the Agent's "Personality"
One level deeper: an agent's memory system decides who it "is." Two agents on the same model — one remembers all your preferences and history, the other starts from zero every time — and users will call the former ten times "smarter," though the model is identical. The 2026 agent-app competition is shifting from "model wars" to "memory wars": whoever's memory is more accurate, cleaner, and better at trade-offs builds the product that feels more human.
But the coin has another side: responsibility. Remembering users obligates you to remember correctly, securely, and deletably. GDPR's "right to be forgotten" becomes brutally concrete in the agent era: when a user says "forget what I just said," can your three layers actually delete cleanly? Design the deletion path on day one — it's a compliance requirement and a trust baseline.
Concretely, deletion must handle each layer: working memory dies with the session, but make sure logs and traces don't keep plaintext; episodic memory gets hard-deleted by user ID, including the corresponding vector-index entries (deleting the DB row but not the index entry is a classic bug); semantic memory is the hard one — rules distilled from multiple episodic memories may "contain" that user's information, and strict deletion requires re-running distillation. The pragmatic engineering compromise: tag every semantic memory with its "contributors," and let deletion requests trigger one incremental re-distillation. It's complex, but it's what "the right to be forgotten" really means in the agent era — not deleting a record, but deleting an influence.
One more 2026 trend: memory is moving from "one agent's private property" to "a shared layer across agents." Can preferences a user taught one agent be inherited by their second and third agents? Whoever ships usable "memory sync protocols" first defines the entry point of the next agent ecosystem. That's also why memory should be designed as an independent service now, not coupled inside some agent — architectural decoupling is what qualifies you to play that game later.
Before building, answer three questions: What does my agent need to remember? What's the cost of misremembering? What's the mechanism for forgetting? If you can't answer the third, don't build a memory system yet. A good memory system makes an agent feel like a reliable old friend; a bad one makes it a grudge-holding, forgetful mess — and the entire difference lies in how "forgetting" is designed.
Related articles

On October 3, engineer Kevin Liao published a polemic that hit the HN front page: agent memory plugins are a lottery over RAG snippets; what agents need is a documentation workspace. The essay's diagnosis, its open-source Operator Memory plugin, the two strongest objections, and the minimal practice you can start tonight.

Gergely Orosz visited OpenAI, Anthropic, Cursor, and Ramp and wrote up the 2026 state of the industry: near-100% AI-generated code, agent PRs up ~10x in eight months, code review degrading into theater, the IDE declared legacy. Key takeaways plus three verdicts and four actions for vibe coders.

Traffic is moving from the search box to the AI answer box. A vibe coder ships a product in a week — and nobody finds it. This guide turns the SEO fundamentals (sitemaps, JSON-LD, Core Web Vitals) and the new AI-discovery toolkit (llms.txt, per-page Markdown versions, FAQ schema, agent-readable pricing and API docs) into a shippable 30-day checklist. The core judgment: how well you document sets your product's ceiling in the agent economy.