My AI assistant Ankita can hold long working sessions — research, edits, tests, browser work, all in one go. The problem with long sessions is simple: the model forgets the plan. It starts three things, finishes one, and cheerfully reports that everything is done.
So I built session checklists: a write_todos tool that turns the plan into a small piece of state the agent maintains itself — and, crucially, a piece of state the agent cannot rewrite.
The tool is deliberately strict
The whole feature is one ES module (tools/project/write-todos.mjs). The agent creates a checklist with a todos array, then updates items by stable ID with an updates array. That's it. But the validation rules are where the design lives:
-
Existing items can never be replaced or removed. Try to drop a step and the tool throws:
cannot remove existing todos. Use status: cancelled instead.Statuses arepending,in_progress,completed, andcancelled— cancelled is how you admit a plan changed. It leaves a visible record of scope changes instead of letting the agent silently delete the evidence. -
Reordering is forbidden. If you resubmit the list with different text or a different order, you get
cannot replace or reorder existing todos; use updates with stable IDs.Stable IDs (s1,s2, ...) are assigned on first write, and every item keeps a created timestamp (at) and acompletedAttimestamp. -
Each item has an
activeForm— a present-continuous label like "Writing tests" — so the UI can show what's happening right now in human terms.
The rendered output is deliberately plain:
[x] 1. Research the endpoint (s1)
[>] 2. Write the new tool module (s2)
[ ] 3. Add tests (s3)
The plan comes back into the prompt every turn
A checklist that only exists when the agent remembers to look at it is useless. So refreshPrompt() in the agent core injects the current checklist into the system prompt on every single turn: id, content, and status for each item, plus one line that does most of the behavioural work:
"Before your final reply, use write_todos to mark work actually completed as completed. Leave unfinished work pending or in_progress; do not mark it completed without evidence."
That last sentence — without evidence — is the whole point. The checklist is an anti-hallucination device disguised as a productivity feature. Claiming completed is the only lie that counts, and the prompt says so up front.
Corrective errors, not punitive ones
My favourite design detail is in the error messages. An earlier trace showed the model wasting five write_todos rounds fighting the validation — bad ID, reorder attempt, removal attempt — because each error just said "no". So every error now carries the current valid state:
- Unknown ID → error lists the available IDs (
unknown todo ID: s5. Available IDs: s1, s2, s3...) - Removal attempt → error appends the rendered checklist and reminds it to use
cancelled - Reorder attempt → error appends the rendered list so the model can see the IDs, order, and statuses
The appended rendering is capped at ~2000 chars so a 100-item list can't blow the context window. The expected effect: the 5-round failure becomes a 1-round correction. When a tool is used by a model rather than a human, error messages are UX — they should make the next attempt succeed, not just explain the last failure.
It survives restarts and old threads
Checklists are saved with the session file and restored when the thread reloads. There's also a hidden todo-snapshot tool entry appended to the transcript on load, so the conversation history contains the plan even when the live agent state is gone. And for threads that predate the saved format, a todoProgress() function reconstructs the plan by replaying the thread's completed write_todos calls — re-running todos arrays, applying updates, or even parsing the rendered result text back into items. Your old plans survive the format change.
On desktop, the checklist also gets a visible surface: successful write_todos calls drive a compact, expandable checklist above the composer showing the completion count and the active step, updating live as the agent works. New user turns clear stale plans until the agent updates them, so the UI never shows you a plan for a request you've already moved past.
What it taught me
Session state you can't quietly rewrite beats session state you can. Most agent frameworks give the model a scratchpad; I wanted something closer to an audit log. The difference between "the agent decided not to do step 3" and "the agent silently forgot step 3" turns out to be one enum value.
Ankita is open source: https://github.com/akyourowngames/A.N.K.I.T.A
If you've built plan-tracking into an agent of your own, I'd genuinely like to hear how you handled the "claiming done without evidence" problem — that's the part I'm still not sure I've solved completely.
Top comments (0)