DEV Community

yureki_lab
yureki_lab

Posted on

How I Built Git Checkpoints So I Can Undo Anything My AI Coding Agent Does

TL;DR

My fully autonomous implementation system runs coding tasks while I sleep, and for a long time my only "undo" was git reset --hard plus hope. I replaced that with per-task checkpoints stored on hidden git refs: a snapshot of the entire working tree (untracked files included) before every task, automatic rollback when verification fails, and a restore command that is itself undoable. Here's the ~40 lines of shell that do the heavy lifting, plus the five lessons that changed how I think about agent safety. πŸš€

The Problem

I run an orchestrator module that hands tasks to parallel implementation agents overnight. Most mornings that's great. One morning it wasn't.

The queue had seven tasks. Task 3 was "clean up the config loader." The agent's cleanup was technically passing tests, but it had quietly changed how defaults were resolved. Tasks 4 through 7 were then built on top of that change. By the time I opened my laptop, I had four tasks of perfectly reasonable work stacked on one bad foundation.

My options were all bad:

  • ❌ git reset --hard to before the run: throws away tasks 1–2 (good) and 4–7 (mostly good)
  • ❌ Manually untangle the diff: the "cleanup" was smeared across the same files the later tasks touched
  • ❌ Ask the agent to "undo task 3": it happily produced a new change that approximated the old behavior, which is not the same thing as undoing

And there was a second, sneakier problem. A different night, an agent failed a task halfway, and the next task started from that half-edited working tree. The half-finished edits weren't committed anywhere. They just leaked forward, and the second task got blamed for breakage it didn't cause.

Both failures have the same root cause: I had no cheap, exact, mechanical way to go back. Everything I'd built so far was about preventing bad changes. Nothing was about reversing them.

Claude Code's interactive rewind is nice when I'm sitting at the keyboard, but my agents run headless (claude -p, Claude Code 2.x) under an orchestrator. I needed undo as infrastructure, not as a keypress.

How I Solved It

The design has three rules:

  1. Snapshot before every task β€” the whole working tree, not just committed state
  2. Roll back automatically when a task fails verification, so nothing leaks into the next task
  3. One task = one commit once it passes, so landed work can be reverted individually
flowchart TD
    A[Orchestrator picks task] --> B[Checkpoint: pre]
    B --> C[Agent implements]
    C --> D{Verification passes?}
    D -- yes --> E[Squash into one commit with task ID]
    D -- no, retries left --> C
    D -- no, out of retries --> F[Restore pre checkpoint]
    F --> G[Mark task failed, attach diff of the attempt]
    E --> H[Next task starts from a clean tree]
    G --> H

Checkpoints live on hidden refs, not branches

My first attempt used git stash. Don't. Stash is a single shared stack, and with several agents working in their own worktrees it became a guessing game of which entry belonged to whom.

My second attempt made real "WIP" commits on the working branch. That polluted history and confused the agents, who would read git log and treat checkpoint commits as meaningful context.

What finally worked: build the snapshot commit without touching the branch or the real index, and park it under a custom ref namespace. Plain git log and git branch never show it, but it's a real commit object that git won't garbage-collect while the ref exists.

checkpoint() {
  local task_id="$1" label="$2"
  local tmp_index; tmp_index="$(mktemp -u)"

  # Work on a copy of the index so the agent's staging area is untouched
  cp "$(git rev-parse --git-dir)/index" "$tmp_index"
  GIT_INDEX_FILE="$tmp_index" git add -A

  local tree commit
  tree="$(GIT_INDEX_FILE="$tmp_index" git write-tree)"
  commit="$(git commit-tree "$tree" -p HEAD -m "checkpoint: $task_id ($label)")"

  git update-ref "refs/checkpoints/$task_id/$label" "$commit"
  rm -f "$tmp_index"
}
Enter fullscreen mode Exit fullscreen mode

The trick is GIT_INDEX_FILE. git add -A against a temporary index captures modified and untracked files into a tree, and commit-tree wraps that tree in a commit whose parent is the current HEAD. The branch pointer never moves. The real index never changes. From the agent's point of view, nothing happened.

On a mid-sized repo this takes well under a second, because git only writes blobs for files that actually changed.

Restore is exact, and it's undoable

restore() {
  local task_id="$1" ref="refs/checkpoints/$1/pre"
  git rev-parse --verify --quiet "$ref" >/dev/null \
    || { echo "no checkpoint for $task_id" >&2; return 1; }

  # Undo must be undoable: snapshot the current mess first
  checkpoint "$task_id" "before-restore-$(date +%s)"

  git add -A                                            # make new files visible to git
  git restore --source="$ref" --staged --worktree -- .  # tree now matches the checkpoint
  git reset -q                                          # unstage; leave the files alone
}
Enter fullscreen mode Exit fullscreen mode

Three lines do the real work:

  • git add -A stages everything, so files the agent created during the failed attempt are known to git
  • git restore --source=<ref> --staged --worktree makes the tree match the checkpoint exactly: edits reverted, deleted files back, newly created files gone
  • git reset -q returns the index to HEAD, so files that were untracked before the task are untracked again

Before any of that, restore takes its own checkpoint. The first time I rolled back the wrong task ID, that one line saved me an afternoon.

Wiring it into the orchestrator

The orchestrator loop is boring on purpose:

run_task() {
  local task_id="$1"
  checkpoint "$task_id" pre

  for attempt in 1 2 3; do
    run_agent "$task_id" "$attempt"
    if verify "$task_id"; then
      git add -A
      git commit -q -m "$(task_title "$task_id")" -m "Task-Id: $task_id"
      return 0
    fi
  done

  # Keep the evidence, then go back to a clean slate
  checkpoint "$task_id" failed
  git diff "refs/checkpoints/$task_id/pre" "refs/checkpoints/$task_id/failed" \
    > "reports/$task_id.failed.diff"
  restore "$task_id"
  mark_failed "$task_id"
}
Enter fullscreen mode Exit fullscreen mode

Two details matter here.

The failed attempt is preserved, not destroyed. The failed checkpoint and the saved diff mean I can inspect exactly what the agent tried, or even resume from it by hand. Rolling back the working tree is not the same as throwing away information.

Every landed task carries a Task-Id trailer and is a single commit. That's what fixes my original seven-task disaster. Finding and reverting one task is now mechanical:

git revert "$(git log --format=%H --grep="Task-Id: task-003" -1)"
Enter fullscreen mode Exit fullscreen mode

If the revert conflicts with later tasks, at least the conflict is honest and localized, instead of me reverse-engineering a smeared diff at 7 a.m.

Housekeeping

Checkpoint refs keep objects alive forever, so I prune anything older than 14 days from a nightly job:

cutoff=$(( $(date +%s) - 14 * 86400 ))
git for-each-ref --format='%(refname) %(committerdate:unix)' refs/checkpoints |
while read -r ref ts; do
  [ "$ts" -lt "$cutoff" ] && git update-ref -d "$ref"
done
Enter fullscreen mode Exit fullscreen mode

Custom refs also aren't pushed by default, which is exactly what I want. They're local scaffolding, not history.

What checkpoints don't cover ⚠️

I tested this by deliberately trashing a scratch repo, and the gap showed up immediately: gitignored files are not restored. git add -A skips them by design. That means .env files, local SQLite databases, and build caches are outside the safety net.

So I split state into three buckets:

State Undo mechanism
Tracked + untracked source files git checkpoint (above)
Ignored-but-important local state (dev DB, env files) plain file copy into a per-task snapshot directory before the task
Anything outside the machine (deploys, emails, third-party API writes, pushed commits) no undo exists β€” requires human approval up front

That third row is the important one. A checkpoint system makes you feel invincible, and that's precisely when an agent runs a migration against a shared database.

Lessons Learned

1. Cheap undo beats perfect prevention.
I spent months adding guardrails to stop agents from making bad changes. Guardrails are worth having, but they're probabilistic. Undo is deterministic. Once rollback cost dropped to one command, I could loosen several rules that had been slowing the agents down, because the worst case went from "lost morning" to "lost task."

2. Snapshot the working tree, not just commits.
"It's in git" is only true for committed state. Agents live in the uncommitted zone: half-written files, generated fixtures, scratch scripts. If your checkpoint ignores untracked files, your restore will leave debris behind and the next task inherits it.

3. Never ask the agent to undo its own work.
An LLM "undoing" a change is writing a new change that it believes is the inverse. Sometimes it is. Undo should be a mechanical operation with zero model calls in the loop.

4. Make reversibility a property of how work lands.
One task, one commit, one machine-readable ID. This costs nothing at commit time and it's the difference between git revert and archaeology. If I can't revert a task in one command, the task was too big or the landing process was wrong.

5. Sort every action by reversibility before it runs.
The useful question isn't "is this action dangerous?" but "can I take it back?" File edits: yes, freely allowed. Local database changes: yes, with a file snapshot. External side effects: no, so they wait for a human. That one axis simplified my permission rules more than any allowlist did.

What's Next

  • Mid-task checkpoints. Right now I snapshot per task. For long tasks I want a checkpoint after each passing verification step, so a failed attempt can resume from the last good state instead of from zero.
  • Smarter retries. Attempt 2 currently continues from attempt 1's working tree. I'm experimenting with restoring to pre first and passing the failed diff as context instead ("here's what didn't work"), which looks promising for tasks where the first approach was a dead end.
  • Restore from the remote control dashboard. Rolling back a task from my phone should be one tap, with the "before-restore" checkpoint making that tap safe.

Wrap-up

If you run coding agents unattended, build the undo button before you build the tenth guardrail. The whole thing is a temporary index, commit-tree, and a ref namespace: about 40 lines of shell on stock git (I'm on git 2.47, but nothing here needs anything newer than git restore, which arrived in 2.23).

πŸ’‘ Try it this week: add the checkpoint function to whatever script launches your agent, and run it before every task. You'll forget it exists until the morning you need it.

If this was useful, follow me here on Dev.to. I write regularly about what actually breaks when you let AI agents ship code on their own, and how I fix it. And I'd love to hear in the comments: what's your undo strategy for agent-written changes? πŸ‘‡

Top comments (1)

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

You reinvented refs/stash in the best way β€” stash is one ref with a reflog of snapshot commits, and your namespace is the same shape generalized to N checkpoints, which is the part git never shipped. Two questions from running agents against git daily: do the checkpoint refs ever expire, or does refs/checkpoints grow unbounded across months of tasks? And does the agent get diff access to them β€” checkpoint N as a readable snapshot is what turns undo into an audit trail, not just a fire exit.