Back to Explore
GuideVibeFix 编辑部Updated Oct 7, 2026

When Your Agent Dies, Don't Start Over: Three Checkpoint Layers for Resumable Long Tasks

Long tasks die three ways: context explosion, process death, or human kill — all voiding your progress. This guide gives you three checkpoint layers: commit discipline for code, data snapshots for data, phase summaries for process — plus idempotent design and a monthly 10-minute recovery drill. Ctrl+C becomes a hiccup, not a disaster.

Close-up of code: minified JavaScript on a dark background

How long tasks die

Ask every vibe coder: what was your longest agent run? And — how did it die? The answers fall into three kinds of death: context explosion (the conversation got too long and the agent started hallucinating), process death (terminal closed, network dropped, machine slept), or human kill (it looked wrong, so Ctrl+C).

All three share one trait: all prior progress is voided. Two hours of research, thirty minutes of dependency installs, an hour of test runs — one Ctrl+C and you're back to square one. Earendil's just-released Pi Durable (which we covered yesterday) productizes "checkpoint and resume" as agent execution infrastructure — but you don't need to wait for frameworks to mature. Three layers of checkpoints arm your long tasks today.

Git branch diagram: feature branches forked off Master, independently revertable

Layer 1: commit discipline — the cheapest checkpoint

One rule: commit after every independently verifiable step. Not "commit when the whole feature is done," but "migration written and tested → commit; API endpoint added and self-tested → commit; frontend wired and page runs → commit."

Put it in your AGENTS.md or project instructions: "git commit after each subtask, with the commit message stating how it was verified." Agents are obedient — tell it to commit per step and it will. Then when the task dies at step 8, steps 1–7 are lying safely in git history. Recovery isn't "start over," it's "continue from step 8."

Advanced move: have the agent branch before starting. Long tasks always get a feature branch; if it dies, drifts, or breaks things, delete the branch and main stays pristine. More reliable than any "undo prompt," because git doesn't hallucinate.

Layer 2: data snapshots — insurance for the irreversible

Code rolls back with git; data doesn't. Stories of agents dropping databases, botching migrations, or corrupting production data surface every month. Checkpoint thinking lands here as: before any agent touches a database, have a rollback-able snapshot.

In order of cost: branch features on platforms like Supabase/Neon (one command forks a full data copy; the agent plays in the branch, merges back after verification); scheduled snapshots plus a manual one before tasks on traditional databases; at minimum, down migrations for every migration script — if up runs, down must roll back.

And one iron rule: agents never connect directly to production. Staging or a branch copy is their playground; production write access belongs only to migration scripts you've reviewed. The cost is one extra deploy; the payoff is one fewer resume-generating event.

Layer 3: runtime checkpoints — borrowing from durable execution

The first two layers protect output; the third protects process. The core ideas of durable-execution frameworks like Pi Durable, Temporal, and Restate distill into three practices you can use without any framework:

1. Steps as records. Have the agent write each step's inputs and outputs to a structured task log (a JSONL file works): step number, tool called, result, duration. When the task dies, you open the log and see "died at step N" instead of trawling thousands of lines of conversation.

2. Design operations to be idempotent. This is what makes checkpoints "resumable" rather than "re-doable": running the same command twice changes nothing. Check existence before creating ("skip if table exists"), overwrite files instead of appending, use idempotency keys on API calls. When agents write idempotent scripts, resuming never worries about "step 3 ran again and duplicated the data."

3. Recovery points are declared, not discovered. Before a long task starts, agree with the agent: "after each phase, output a status summary (done / in-progress / todo / blockers)." That summary is your recovery point — open a fresh session, paste it in, and the agent picks up where it left off. Context windows explode; summaries don't.

Recovery drills: don't wait for a real outage

Like backups, untested checkpoints don't exist. Once a month, spend 10 minutes: deliberately kill a running long task, find the last good commit with git log, branch from there and continue; open a fresh session with a phase summary and see if the agent picks up seamlessly; restore a database branch snapshot and confirm the flow works. Any "stuck point" the drill reveals — a thin summary, an unmergeable branch — is the cheapest tuition you'll pay before a real incident.

Three anti-patterns

Anti-pattern 1: checkpoints too coarse. "Commit when done" is no checkpoint at all. The test: kill the task at any moment — you should lose at most 15 minutes of work. More than that, and your commit granularity is too coarse.

Anti-pattern 2: backing up code but not data. The most complete git history won't save a database the agent mangled. Code and data are two separate checkpoint systems; you need both.

Anti-pattern 3: decorative checkpoints. Commits without verification notes, branches never actually verified on, summaries too thin to resume from — checkpoint quality determines recovery quality. After every recovery, ask: was that smooth? Wherever it wasn't is what you fix next time.

The one-line summary

Long-task reliability doesn't come from the luck of "one clean run" — it comes from designing for "resume anytime." Commit discipline protects code, data snapshots protect data, phase summaries protect process. With three layers in place, Ctrl+C downgrades from "disaster" to "hiccup." Frameworks will evolve; the idea won't expire.

Browse projectsPublish your project

Related articles

PromptGit concept art visualizing prompt version control
Guide
Treat Prompts Like Code: Prompt Version Control for Vibe Projects

Prompts in vibe projects live in code strings, admin text boxes, and docs — changed live, version unknown when things break. This guide shows how to treat prompts like code: a prompts/ layout, YAML frontmatter, semantic versioning, PR reviews, canary rollouts with one-click rollback, plus an evals baseline — and a real war story: one added sentence cost 12 points of classification accuracy.

AI CodingDeveloper WorkflowTool Tips
Pull request workflow illustration: a developer submits code while code windows pass check marks toward merge
Guide
After the AI Writes the Code: A Practical Code Review Workflow for Vibe Projects

The faster AI writes code, the more review matters. Four layers: diffs for logic (boundaries, errors, concurrency — plus auth, payments, SQL, encryption, secrets), runtime for behavior (type checks, lint, security scans go green first), AI for first-pass screening (a second model reviews, humans read only flagged parts), humans for the final call (AI never clicks merge). Includes commit norms, PR template, branch protection, rollback plans.

AI CodingDeveloper WorkflowTesting & Quality
Developer team collaborating on code
News
GitHub Was Built for Humans: Cloudflare Offers $25,000 in Credits to Rebuild Git for Agents

Cloudflare's Birthday Week blog makes the case plainly: GitHub was designed for humans writing code; the agent era needs the collaboration layer reinvented. Artifacts enters open beta with a repo for every agent, plus a developer competition — $25,000 in credits for first place, deadline October 14. This is the first time a major infra vendor has put 'infrastructure for agents writing code' on the table as a public proposition.

AI CodingDeveloper WorkflowIndustry Trends