Context Engineering in Practice: It's Budget Management
Context windows keep growing in 2026, but "stuff everything in" buys sky-high bills and diluted attention. This hands-on guide treats context engineering as managing three ledgers: token money (cache everything cacheable, retrieve instead of stuffing, compress history), model attention (important at both ends, structured expression, context decay, reference over copy), and your time (10-minute cap for one-offs, build assets for recurring tasks). Includes a 5-minute context checklist template.

In 2026, context windows keep growing: 1M-token models are unremarkable, 10M is on the way. Many conclude: "windows are huge — just stuff everything in, right?" Then the bill arrives: hundreds of thousands of tokens per task, output no better than "5 carefully chosen files."
This is a hands-on guide to context engineering. The thesis: context engineering is essentially budget management — you keep three ledgers: token money, model attention, and your time. Manage all three and agent effectiveness doubles; mismanage them and the biggest window is waste.
Ledger 1: token money
Do the math. Take the standard H2 2026 price of $2/$10 per 1M tokens (see art9/art11 in this batch). A medium task: 100K input tokens (history + retrieved code), 10K output. Per-run cost = 0.1×$2 + 0.01×$10 = $0.30. Trivial? 100 tasks a day means $900/month. For an indie developer, that's real money.
Three savings levers, ranked by ROI:
- Cache everything cacheable. System prompts, AGENTS.md, project-structure docs — identical every run — go through prompt caching ($0.10–$0.20 per 1M, see art11). Cost drops 90%. The most certain savings of 2026; not using it is leaving money on the table.
- Retrieve instead of stuffing. Don't dump all of src/. Use code indexing/RAG to retrieve the 5–10 most relevant files. In practice, precise retrieval uses ~1/10th the input of "stuff everything" — and works better, because there's less noise.
- Compress conversation history regularly. In long tasks, history is a token black hole. Every 10–20 turns, have the agent summarize ("compress progress, key decisions, and todos into 500 words") and replace the original with the summary. Claude Code's compaction and Codex's auto-compression do this; manual works too.
Ledger 2: model attention
This ledger matters more than money. Model attention is finite — the longer the context, the more diluted the effective attention. "Needle in a haystack" tests proved long ago: where information sits in a long context significantly affects whether the model uses it. A 1M window ≠ a 1M effective window.
Four rules for managing attention:
- Put the important stuff at both ends. Beginnings and endings are attention highlands. System instructions at the start, the current task at the end, reference material in the middle. That's transformer architecture, not mysticism.
- Structured beats natural language. For code structure, show "file tree + interface signatures," not "paste 10 files in full." For requirements, use "numbered lists," not "a prose paragraph." Structured information has higher signal-to-noise; models process it faster and better.
- Give each step only what it needs. At step 5, step 1's detailed logs are useless — a summary suffices. Call it "context decay": information depreciates over time; prune actively. Good agent frameworks do it automatically (compaction); in manual workflows, periodically tell the agent "forget earlier details, keep only conclusions."
- "Reference" instead of "copy." "Follow the error-handling pattern in src/auth/login.ts" beats pasting login.ts in full. Modern agent tools support file references (@, #) — let the model read on demand instead of you deciding what it reads.
Ledger 3: your time
The most overlooked ledger. You spend 30 minutes crafting context to save the agent 5 minutes of trial and error — good trade? Depends on your hourly rate and the task's repeatability.
Decision framework: one-off tasks — max 10 minutes of context prep. Hand over a few key files; let the agent explore the rest (it has tools). Recurring tasks — worth 2 hours building "context assets": AGENTS.md, few-shot example libraries, common retrieval templates. These assets save time every run; ROI compounds.
One anti-pattern: "context perfectionism." Some spend an hour tuning prompts and context for one-shot agent success. But an agent trial-and-error round costs 30 seconds and $0.30 — how many $0.30s is your hour worth? Accept "70-point context + 2 feedback rounds" over chasing "100-point context + 0 rounds." Remember guide7's conclusion: feedback loops are 2026's core skill. Don't spend everything on the first frame.
Working template: a task's context checklist
For each new task, prep context with this checklist — five minutes:
- □ AGENTS.md (project-level; let the agent read it — don't paste it)
- □ Task description: 3 sentences — what, acceptance criteria, constraints
- □ Key file references: 3–5, via @ or #, never full text
- □ Counter-example / example: one "don't do it like this" beats ten "be careful"s
- □ Output format: diff or full files? Explanation or code only? Say it — don't make the agent guess
The three ledgers, quantified: one task's true cost
Enough theory — let's do real math. Say you're having an agent refactor a 5,000-line module:
Plan A: stuff everything in. Input: the whole module ≈ 60K tokens plus 20K of conversation history = 80K. At $2/$10: input $0.16, output 8K tokens $0.08, $0.24 per run. Cheap? But "stuff everything" is noisy — the agent averages 3 rounds to get it right (round 1 misreads file relationships, round 2 fixes, round 3 passes tests). Total $0.72, 25 minutes including your review time.
Plan B: retrieve + cache. Input: AGENTS.md (cache hit at $0.10) + 6 retrieved key files at 12K tokens + 1K task description. Input cost: cached portion at $0.10, 13K new tokens at $2 ≈ $0.027. Output 6K (sharper input → shorter output) $0.06. $0.087 per run, done in one round (because everything given was signal). Total $0.087, 8 minutes.
8x the cost difference, 3x the time. And Plan B's extra cost was just your upfront 10 minutes writing AGENTS.md and tuning retrieval — 10 minutes that are an "asset," reused next task. Do this math once and you'll never "stuff everything" again.
Tool comparison: context management across major agent tools
In H2 2026, mainstream tools' context management falls in three tiers:
- Tier 1 (automatic): Claude Code. Auto-compaction (history compression), tunable context policies in /config, native AGENTS.md support. Least for you to do; fits "don't want to think about it" people. Price: "black box" — you can't see how it compresses, and it occasionally drops key info.
- Tier 2 (semi-auto): Codex CLI, Cursor. @ references, file-level caching, manual compress commands — but "when to compress, what to compress" is your call. Most flexible; fits hands-on intermediate users. Cursor's Rules + @ combo is currently the ceiling of "manual context management."
- Tier 3 (bare metal): raw API calls. Calling APIs directly means your context management is all hand-written: retrieval, chunking, compression, caching, every line yours. Most work, most control — AI product builders live here, because "context strategy" is itself product moat.
Advice: individuals take tier 1 or 2 — don't burn time on tier 3 (unless building a product). Building an AI product requires tier 3, because "managing context well" is your moat.
Context assets, directory layout: a copyable template
Finally, a "context assets" directory layout — steal it:
- AGENTS.md (repo root): the constitution, 100 lines — see art14 in this batch.
- .agent/examples/: few-shot example library. One file per frequent task: input (task description) + output (a good result). The agent reads same-category examples before new tasks.
- .agent/patterns.md: code patterns doc. "How error handling is written here," "what API responses look like" — write conventions down; stop making the agent guess every time.
- .agent/review-checklist.md: the review checklist (guide9 in this batch), for pre-submit self-checks.
- .agent/context-budget.md: per-task-type "context budgets" — "refactor tasks: input ≤ 30K tokens," "bug fixes: error + related files first, ≤ 10K." Over budget means the task needs finer splitting — go split it.
Building this takes 2–3 hours, but it's "invest once, dividends forever." Three months later your agent task success rate, average token cost, and your prep time all improve at once — that's the compound interest of context engineering.
3 anti-patterns: don't manage context like this
Finally, three common "mismanagement" anti-patterns for self-checking:
Anti-pattern 1: "hoarding" — reluctant to delete anything. 200 rounds of history kept "in case it's needed later." Truth: information from 200 rounds ago has <1% chance of future use, yet every round burns money and dilutes attention. Cure: one iron rule — "history beyond 20 rounds must be compressed into a summary." Don't mourn the loss; what compression drops matters far less than what it saves.
Anti-pattern 2: "helicopter feeding" — deciding everything for the agent. Some find all 30 files, sort them, summarize them, then feed the agent. Those 2 hours of "prep" reach 80 points via the agent's own 5-minute retrieval. Your time costs far more than tokens — context prep caps at 10 minutes (one-off tasks); beyond that you're working for the agent.
Anti-pattern 3: "mystical tuning" — guessing retrieval counts. "10 retrieved files underperformed; try 20?" — don't guess, measure. Pin 5 representative tasks, set retrieval counts to 5/10/20, log success rates and token costs, plot the curve. Most people's sweet spot sits at 5–10; past 10, diminishing returns. Data over gut — that's what "engineering" means.
Tool picks: where to start
Unsure where to begin? Try in this order: individuals start with Claude Code's auto-compression (zero config); if the "black box" worries you, switch to Codex CLI or Cursor's manual @ references. Building an AI product? Go straight to LangChain's retrieval + compression components — don't hand-roll. The principle: automatic first, manual when you hit the ceiling — automatic tools already solve 80% of most people's context problems.
The bottom line: context engineering isn't "how to stuff more" — it's "how to spend less and achieve more." Three ledgers — token money, model attention, your time — each deserves careful accounting. Bigger windows are good news, but they reward not "who stuffs most" but "who manages best." In H2 2026, vibe coding competition has shifted from "whose model is stronger" to "whose context management is better." And the latter costs no money — only thought. That's the fairest battlefield an indie developer could ask for.
Related articles

On October 3, engineer Kevin Liao published a polemic that hit the HN front page: agent memory plugins are a lottery over RAG snippets; what agents need is a documentation workspace. The essay's diagnosis, its open-source Operator Memory plugin, the two strongest objections, and the minimal practice you can start tonight.

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

Gergely Orosz visited OpenAI, Anthropic, Cursor, and Ramp and wrote up the 2026 state of the industry: near-100% AI-generated code, agent PRs up ~10x in eight months, code review degrading into theater, the IDE declared legacy. Key takeaways plus three verdicts and four actions for vibe coders.