Context Engineering: Teach Your Agent What to Remember and What to Forget
When an agent drifts off course mid-task, the context window is usually the culprit. This guide covers practical memory management: what belongs in the system prompt, what belongs in files, when to clear context — and three techniques for long tasks that don't derail.

Why agents get 'dumber' the longer they run
You've seen it: an agent starts strong, then after a few thousand lines starts acting confused — re-asking things you already said, breaking B while fixing A, forgetting the project's core conventions. The problem isn't the model; it's the context. An agent's working memory is a finite window: system prompt, conversation history, files it has read, tool outputs — all crammed in. When the window fills with irrelevant material, the actually important conventions get pushed to the edge of attention, and the model starts acting on fuzzy impressions.
Context engineering is the craft of actively managing what the agent sees, in what order, and when it forgets. The difference from prompt engineering: prompting is "how you say it," context engineering is "what you let it see."
Three layers of memory: put things where they belong
In practice, split information into three layers, each with its home. Layer one is the 'constitution' — project conventions that never change (code style, directory layout, prohibitions). Put them in files the agent always reads at startup, like AGENTS.md or CLAUDE.md. Write once, lasts forever. Layer two is the 'task brief' — background relevant to this task only (the requirement, the three files involved, known pitfalls). Put it in the prompt you send; discard when the task ends. Layer three is 'working notes' — intermediate conclusions produced mid-task ("tried approach A, failed because of X"). Have the agent write these to a scratch file (e.g. .agent-notes.md) so it doesn't repeat mistakes after forgetting.
The most common mistake is stuffing everything into the prompt: the longer the prompt, the more diluted the critical information becomes. One rule of thumb: the prompt carries only what this task cannot proceed without; long-lived conventions live in files.
Three techniques for long tasks that stay on track
First, do deliberate 'shift handoffs.' When the agent starts drifting mid-task, don't keep chatting — have it stop and summarize in a file: "where we are, what's next, known pitfalls," then open a fresh session that reads the summary. Far more effective than wrestling inside a tens-of-thousands-of-tokens conversation.
Second, use files instead of chat history. Have the agent write decisions and conclusions into markdown in the repo (design decisions, API contracts, todo progress) instead of leaving them in the chat log. Files are searchable, version-controlled memory; chat history is not.
Third, 'compact' regularly instead of wiping. When context fills up, don't just start a new session and lose everything — have the agent distill a state summary first (done, in progress, pending, key constraints), and carry it into the new session. That's what compaction does in production agent systems; you can do it manually too.
The one-line summary
Treat the agent like an intern with limited memory but strong execution: write important things into files (long-term memory), state the current task clearly (short-term memory), and persist conclusions as you go (the notebook). Good context management can double an agent's effective intelligence; without it, even the strongest model can't complete long tasks reliably.
Related articles

Prompts in vibe projects live in code strings, admin text boxes, and docs — changed live, version unknown when things break. This guide shows how to treat prompts like code: a prompts/ layout, YAML frontmatter, semantic versioning, PR reviews, canary rollouts with one-click rollback, plus an evals baseline — and a real war story: one added sentence cost 12 points of classification accuracy.

The faster AI writes code, the more review matters. Four layers: diffs for logic (boundaries, errors, concurrency — plus auth, payments, SQL, encryption, secrets), runtime for behavior (type checks, lint, security scans go green first), AI for first-pass screening (a second model reviews, humans read only flagged parts), humans for the final call (AI never clicks merge). Includes commit norms, PR template, branch protection, rollback plans.

Getting signups is only the start — users churn by day 3 and you have no horn to call them back. This guide covers notification systems for vibe projects: channel selection, email with Resend from day one, SPF/DKIM/DMARC done right, when SMS is worth the money, frequency caps and unsubscribe, retries and dead letters, plus a launch acceptance checklist.