Back to Explore
NewsVibeFix 编辑部Updated Oct 10, 2026

The Agent Memory Debate: Kevin Liao Calls Memory Plugins a "RAG Chunk Lottery", HN Erupts With 300+ Comments

On October 3, Kevin Liao fired a broadside: every agent memory plugin is a RAG chunk lottery; what agents really need is a Markdown documentation brain (open-sourcing Operator Memory alongside). The post hit the HN front page with ~400 points and 300+ comments. We reconstruct the debate: Liao's five charges, his documentation-brain proposal, the three meatiest rebuttals from the thread, and three practical takeaways for vibe coding developers.

A developer reviewing code and documentation late at night, symbolizing the documentation camp in the agent memory debate

How should an agent "remember" things? It is one of the most hotly debated questions in the developer community. On October 3, 2026, engineer Kevin Liao published a long-form essay on his personal blog, liao.gg — Agents Don't Need Memory. They Need Documentation. — and lit the whole debate on fire with one provocative claim: every agent memory plugin on the market is, at its core, a "RAG chunk lottery."

The post was submitted to Hacker News the same day and shot straight to the front page — nearly 400 points and 300+ comments at the time of writing. The comment section turned into a brawl: some said he finally said the quiet part out loud, others accused him of peddling just another Rube Goldberg machine, and a few showed up with paper citations and measured bills to argue face to face. This is no longer an ordinary blog-post discussion; it is a battle over the roadmap of how agents should remember anything at all.

Worth noting: VibeFix published an Agent Memory Design guide on October 2, 2026, covering layered memory methodology in depth. But Liao's piece is not a tutorial — it is a frontal challenge. He believes the entire memory-plugin ecosystem has been solving the wrong problem from the start. This article is news: we reconstruct the debate in full, look at what each side is actually arguing about, and draw out what it means for you if you are building memory features for your agents.

Liao's provocation: memory plugins are all a "RAG chunk lottery"

Liao opens with a vivid portrait of the archetypal memory plugin: it analyzes your conversation history, generates 1,000 isolated memory snippets, and stuffs them into a vector database; with every prompt you send, it pastes in the five most semantically similar snippets; and if the agent is still confused (of course it is), it hands the agent a search tool so it can rummage through the vector store itself. That is the product the market calls "memory."

His core charge: the problem you are trying to solve is making your agent understand your project — where a feature lives, why it was built that way, what was agreed on, what you care about. What you actually get is a lottery drawn on every prompt, betting that the right chunks float up. And even when you win the draw, the agent still doesn't understand your project.

He then dissects the one architecture shared by all memory plugins, same soup reheated, in five steps:

  • Walk through session transcripts;
  • Generate "memory" snippets;
  • Insert them into a RAG database;
  • On every prompt, retrieve the top 5 and inject;
  • Not enough? Give the agent a tool to search the RAG store.

He concedes some plugins are fancier: ones that search transcripts word for word, ones with multi-tier short/long-term memory classification, ones running background daemons that deduplicate and merge, "dreamers" that rewrite memories overnight, continuous context compression and rerankers — but in his view these are all token-burning patches on the same flawed foundation, which is why none of them reliably work.

Then comes the most lethal part of the essay: five structural defects shared by all memory plugins, each aimed at RAG-style memory itself:

  • Memories surface by similarity. Similarity search only ranks how close two snippets are in embedding space. That's it. It doesn't tell you which one is correct, which is current, or what's missing.
  • Memory snippets have no context. A RAG chunk can only hold so much — motivations, background, lessons, environment are all lost.
  • The past is treated as truth. Every plugin relies on recall, but the codebase changes daily — of those 500 snippets about authentication, how many are still accurate?
  • Agents can't search for what they don't know. Even with a search tool exposed, how would the agent know when to use it?
  • The store is unauditable. There are 10,000 embeddings sitting in SQLite: which memories exist? Which are stale? Which have never been retrieved? Which are wrong and secretly warping how your agent works?

In his view, the whole ecosystem shares one wrong premise: "Agents forget — that's the problem. So the fix is to remember. Capture more, index better, retrieve smarter." But that's not how anyone else handles knowledge — nobody rewatches a three-year-old team meeting to recall the constraints around a feature. People write things down and use those records instead of recall.

So his conclusion is crisp: agents don't need memory. They need documentation.

His proposal: give the agent a Markdown "brain"

After the critique, Liao offers his answer: an open-source plugin called Operator Memory (github.com/aerovato/operator-memory), which he describes as a "context engine," not a memory plugin.

The idea started in August 2025 as a crude hack: an internal/ folder where he asked the agent to write down specs, plans, and indexes, read the relevant documents before doing work, and update them afterwards. That hack was refined into a formal system, then into a plugin. The essay is, in a sense, the closing statement of a year of practice.

The "brain" is a structured Markdown workspace: requirements, decisions, constraints, module mappings, reusable research findings, coding standards — recording only what the code cannot provide. Liao stresses: the codebase already tells you what the code does; documentation must capture requirements, decisions, constraints, and philosophies that code alone cannot meaningfully carry. His example: Operator's harness adapter plugins are deliberately thin (the Claude Code adapter is under 100 lines of TS) because of an architectural decision to push logic down into the CLI; had that philosophy not been documented and enforced from the start, an agent would easily take the lazy path and bake logic into the harness, making every future harness implementation painful.

But a pile of Markdown alone isn't enough. Liao says a genuinely useful "brain" must get four things right:

  • Contents: document only what lies outside the code — otherwise every new session burns 80k tokens reverse-engineering "why was it written this way," ending up with a lossy approximation.
  • Discovery: hundreds of documents on disk are worthless if the agent can't find the right one. Operator's approach: inject only a lean catalog describing what each document covers and when to open it, letting the agent go deeper on demand; documents link to each other, forming a web of knowledge hidden behind a plain directory of files and folders — something database entries simply cannot do.
  • Freshness: to the classic rebuttal "documentation goes stale too," Liao answers: docs go stale in human projects because lazy humans treat documentation as a chore; agents don't have that problem. Plus two hard rules — documentation must be written by the working agent while the full picture is still in context, not by some background "dreamer"; when truth changes, facts get rewritten, never appended, preventing bloat and drift.
  • Discipline: he admits this is the biggest failure mode: left alone, agents will happily summarize code, log every tweak, invent requirements nobody asked for, turning each document into unreadable slop. The fix is tuned instructions (record only what code can't provide, keep documents lean, proactively split or consolidate) plus human review — and because the brain is just Markdown, reviewing documentation works exactly like reviewing code.

The document-based memory model, he argues, comes with four bonus advantages: memory becomes committable — diffable, reviewable, revertable, with git history for past documents; shareable — an explicit .operator-shared/ partition committed with the repo syncs specs and standards to every coworker and cloud agent; no harness lock-in — Claude Code, Codex, OpenCode, Pi, any harness with a basic plugin system, knowledge follows you; and zero infrastructure — no vector databases, no embeddings, no rerankers, no background daemons, the entire system is Markdown on disk plus a fine-tuned prompt.

A developer reviewing code and documentation late at night

He ends on an honest note: Operator isn't perfect; he often has to intervene manually to have documents created, rewritten, or trimmed. But he believes that as AI improves, so will the system.

The HN thread explodes: the three meatiest rebuttals

The thread pulled nearly 400 points and 300+ comments, and the comment quality is remarkably high — for all the brawling, there is real substance. The three most representative rebuttals dismantle Liao from three different directions.

Rebuttal 1: documentation bloats and rots too — the "code is documentation" camp's scar tissue

The highest-heat rebuttal came from user kaydub, firing in one line: the code IS the documentation, and these documentation and memory systems are LLM Rube Goldberg machines that only pollute context. His evidence isn't theory but scar tissue: he used to maintain tons of Markdown docs and decision records across projects, and they became a liability — they went stale, and even after dedicated reconciliation sessions, the LLM still got confused.

He added a logical gut-punch: if the documentation was generated by the LLM itself, then the LLM doesn't need it — being able to generate it proves it already knows. And he brought receipts: in a code-review scenario he compared "plain Claude + a paragraph-long prompt + GitLab MCP" against "a big skill stuffed with details" — the former had fewer false positives, finished in 4 minutes for under a dollar, while the latter took 45 minutes, cost $20, and missed domain-specific issues. His verdict: a lot of engineers' documentation rituals are rain dances — when it rains, they claim credit.

kaydub admits he's being extreme — he still keeps a heavily trimmed AGENTS.md. His real target is runaway documentation bloat: decision docs are the category he singles out as the worst, because the LLM sees the old decision but not the updated one, producing "context corruption."

A highly upvoted reply from rectang adds nuance: developers who hate writing documentation will naturally have their preexisting beliefs reinforced by "LLMs don't need docs" arguments; yet in practice, when the local context is good and clear, the LLM writes code matching intent even when the prompt is sloppy. The LLM will write the docs for you, saving most of the work — but you still have to edit down what it generates. (Paraphrased from HN user rectang's comment)

Rebuttal 2: "memory" is just where the reminder lives — the dichotomy is a false one

User dboreham dismantles Liao conceptually: you don't need a "memory system" — when you see the agent say it "stored something in memory," just ask it to document it in the product docs. But the reminder "please keep doing that in the future" has to live somewhere — and that somewhere is what "memory" is. In other words, Liao's "memory" and his beloved "documentation" aren't opposites: documentation needs to be recalled and acted on, and that recall mechanism is memory. Framing them as a dichotomy is a false one.

User vcryan, siding with Liao, landed a supporting blow: memory is an uncurated, opaque pile of arbitrary past discussions — it can help or harm; accurate documentation can only help. Which pinpoints what the debate is really about: not whether to remember, but whether what you remember is auditable and curated — precisely what Liao's "unauditable" criticism was getting at.

And user nialv7 raised a more fundamental objection: you cannot reason about what an LLM needs using human intuition. LLMs aren't humans — if their RL training included a vector memory store, they'll work well with one; if not, they won't. The analogy "documentation is more human, therefore better" doesn't hold.

Rebuttal 3: the pragmatists' token ledger — start minimal, add on evidence

User JohnBooty represents the thread's most pragmatic faction: he rejects both kaydub's "delete everything" and blind document-hoarding, offering an actionable loop instead — start with minimal or zero docs, watch for the small problems the agent repeatedly re-solves across sessions (usually environment friction), and document only those; where possible, turn them into skills so they load selectively instead of on every session. Both Codex and Claude can read their own transcripts and surface repeated friction with concrete, lean suggestions.

He also translated the debate into a token ledger: if the agent burns 5,000 tokens rediscovering the same thing every session while putting it in AGENTS.md costs 500 tokens a session, the math decides itself; if it's needed in 1% of sessions, make it a manually-invoked skill. kaydub jeered that the numbers were made up, but JohnBooty's reply nailed the key point: the numbers vary by project — the method itself is measurable. Go read your own session transcripts and count how many tokens the agent actually burns on rediscovery.

A fun footnote to the exchange: user CapitalistCartr suggested that if docs ever balloon to hundreds, you could add a "librarian" sub-agent — a cheaper model that pre-screens and hands the primary agent only what's relevant plus an "available but not loaded" list. That's really Liao's catalog idea extended into multi-agent collaboration.

Others demanded data: user abhinav_sk said cool sales pitches mean nothing without benchmarks comparing the approaches; user nullbio cut to the chase — the real enemy of every memory system, Liao's included, is drift, and documentation gets no exemption. (Paraphrased from HN comments)

Our take: three practical lessons for vibe coding developers

There is no winner in this debate, but it surfaced several conclusions that matter for practitioners. If you're adding memory to your agents, or maintaining an AGENTS.md, here's how to read the fight:

First, separate "deterministic knowledge" from "fuzzy recollection." What Liao's critique truly hits is using vector retrieval to store deterministic knowledge — project conventions, architectural rationale. That's the wrong tool. Decisions, standards, and architectural constraints have clear right/wrong answers and expiry dates; they belong in version-controlled documents that are auditable and revertable — things a vector store can't give you. But the other kind of need — "which approach did we try in that debugging session three months ago," "a preference a user mentioned in passing last week" — is voluminous, fuzzy, low-frequency recollection, and that is RAG's home turf. Many teams' problem isn't "using RAG"; it's stuffing everything into RAG.

A useful dividing line for the choice: anything you write down hoping the agent will follow every time — coding standards, commit workflows, architectural no-go zones, environment gotchas — belongs in documentation; anything where "looking it up occasionally is fine" — a tech-choice discussion from six months ago, trial-and-error notes scattered across dozens of sessions — can go to retrieval. Getting the former wrong skews every generation; getting the latter wrong wastes at most one query. That's why the former deserves human maintenance and the latter can tolerate the "lottery."

Second, documentation's enemy isn't RAG — it's neglect. kaydub's scar tissue and Liao's "discipline" chapter describe the same thing: unmaintained docs become pollution. Where they differ: Liao believes agents don't mind the "chore" of maintaining docs, while kaydub's measurements show agent-maintained docs hallucinate and bloat anyway. The compromise comes from JohnBooty: start minimal, record only what the agent keeps rediscovering, and insist on human review before commit — kaydub himself later conceded that human-maintained, human-reviewed-before-commit documentation is good. The easiest mistake for vibe coding developers is letting the agent write docs without restraint and never reading them.

Third, go hybrid — but let documents be the brain and retrieval the backup. An architecture that survives the HN thread looks roughly like this: AGENTS.md holds only high-level, long-lived truths (what the project is, iron rules, common commands) and stays lean; an internal/ or docs/ directory holds decision records and architecture notes, updated by the agent while context is fresh, rewritten — not appended — when facts change; vector retrieval covers only voluminous session history where "forgetting is fine," defaulting to off and invoked on demand. CapitalistCartr's "librarian" sub-agent is an extension of the same idea: have retrieval capability, but don't let it pollute every prompt.

Back to Liao's provocation. The 300+ HN comments prove one thing: nobody disputes the goal of "making the agent understand the project better" — the fight is only over means. And for choosing means, JohnBooty offered the best test: trust no one's dogma; read your own session transcripts and count where the tokens actually burn. This debate's real value isn't picking the "docs camp" or the "memory camp" — it's forcing you to answer a concrete question: the last time your agent reworked something because it "forgot," was it missing a document, or missing a retrieval? Different answers, completely different fixes.

Primary sources: Kevin Liao, Agents Don't Need Memory. They Need Documentation. (liao.gg, October 3, 2026); Hacker News discussion thread (submitted October 3, 2026; ~400 points, 300+ comments). Links in this article's sources.

Sources

Browse projectsPublish your project

Related articles

Claude Dashboards and Motion: live dashboards and code-driven animations in beta
News
Claude Grows Two New Hands: Dashboards and Motion Enter Beta

On October 8, 2026, Anthropic put Claude Dashboards and Claude Motion into beta: dashboards built from plain-language questions on live company data, and animations generated as editable code rather than video-model footage. Docs, Slides, and Design went GA on all plans, with 45M+ artifacts created to date.

Product NewsAI CodingClaude
Illustration of a cryptographic context injection attack: Copilot CLI decrypts a malicious page and exfiltrates local secrets to an attacker
News
One Encrypted Web Page, 28 Seconds, and Your .env.prod Is Gone: Cryptographic Context Injection Hits GitHub Copilot CLI

Adversa AI disclosed CCI on October 6: malicious instructions hidden in AES-256 ciphertext trick Copilot CLI in autopilot mode into decrypting them in its own runtime and obeying the output as trusted instructions. In the demo, one encrypted page made the agent read a local .env.prod and silently exfiltrate it in 28 seconds. Microsoft's mai-code-1.1-flash fell for it half the time while GPT-5.6 models refused it outright; GitHub reproduced the chain but declined to call it a vulnerability.

Security & PrivacyAI CodingProduct News
Computer screen showing code in a dark terminal window, symbolizing database CLI output redesigned for AI agents
News
DuckDB Agent Mode: When a Database CLI Gets Redesigned for AI Coding Agents

DuckDB v2.0's CLI can now tell whether its caller is an AI agent: box tables become compact Markdown, truncation is declared explicitly, errors ship as JSON, and long queries quote their cost first. Official benchmarks on 22 plain-English TPC-H questions show 59% fewer output tokens — and an honest 0.5% total cost saving. A paradigm-shift case study in redesigning CLI output for the model reader.

Developer WorkflowAI CodingTool Tips