There is a Grimm fairy tale called The Elves and the Shoemaker. A poor shoemaker cuts his last piece of leather, leaves it on his workbench, and goes to bed. When he wakes up the next morning, a beautifully finished pair of shoes is waiting for him. During the night, a band of elves had crept in and quietly stitched them together.
As a child, I thought it was just a fairy tale.
Lately, though, I've started leaving a single ticket out before bed, typing this into my terminal, and going to sleep:
/graph-ops:autopilot-tree <ticketId>
When I wake up and glance at the screen, the work is done. A plan has been drawn up, the work has passed review after review from several different angles, the tests have been run, and a PR has been opened. Even the bugs found along the way have been filed as separate tickets, and those have been taken care of too.
This article is a sequel to my earlier introduction, "Introducing GraphOps: An Easy-to-Use Graph Engineering Plugin for Claude Code
". That one covered how GraphOps works and how to use it. This time, I want to share what I learned from spending a few nights with the Autopilot feature that has since been added to GraphOps, letting the elves work while I sleep.
A Quick Recap
GraphOps is a ticket management plugin for Claude Code. File a ticket and run /graph-ops:process-ticket, and the work is assembled into an execution graph and carried out step by step.
When I wrote the previous article, a person still had to click an "Approve" button at every approval gate along the graph. Is this plan good to go? Is it OK to release? And so on.
The shoemaker still had to stay up at his workbench.
Enter the Elves
Autopilot is a mechanism that drives a ticket all the way to the end on its own judgment, instead of waiting for a person at those approval gates. It comes in two modes:
-
/graph-ops:autopilot-ticket <ticketId>: carries a single ticket through to the end -
/graph-ops:autopilot-tree <ticketId>: carries that ticket, plus every ticket derived from it, through to the end
You can of course also start it from the "Autopilot" button in the Web UI. The confirmation dialog lets you choose between "Tree (this ticket and its descendants)" and "This ticket only."
And there isn't just one elf at work here.
- The master elf (orchestrator): This is the session where you typed the command. The master elf doesn't sew any shoes. It only plans out the work.
- Lead elves (child sessions): Independent Claude Code sessions that the master elf summons in separate terminals. Each elf takes charge of one ticket and, on a git worktree and branch dedicated to that ticket, handles everything from start to finish: refining the ticket, processing the execution graph, releasing, and deciding what to do with the carry-over items. When it's done, it sends the master elf nothing more than a short, one-to-three-line report.
- Worker elves (subagents): Agents that a lead elf calls up for each node in the execution graph. One elf writes the plan, another implements, another reviews. During parallel reviews, many elves work at the same time.
The approval gates that used to wait for a person are now judged by the elves themselves. They check the deliverables against the ticket's completion criteria and approve unless they can name a concrete defect. Each decision is saved as a deliverable along with its reasoning, so you can read it back in the Web UI once morning comes.
Layer upon Layer of Review Gates
Is it really safe to leave everything to the elves?
That was my first thought, too. But as I watched the execution graphs, I noticed that the elves' work passes through layer upon layer of review gates.
The lineup of review gates is assembled per ticket. Starting from the plan and its plan review, the elves look at what the ticket involves and decide which reviews to include, whether to write tests, and whether to update documentation. There is no fixed template set in advance.
For a ticket that involves implementation and tests, for example, the review gates line up like this:
- A review of the plan, then approval of the plan
- A review of the test cases (the Gherkin test specification)
- Parallel reviews of the implementation, such as code, QA, security, and non-functional reviews (which perspectives are included depends on the ticket)
- A review of the test results
- If the ticket updates documentation, a review of that documentation
- If a report is written, a review of that report
- Finally, approval of the release
The reviews don't just go around in circles forever. With each round, the tier is raised (Normal / Important / Final), and minor findings are passed along as "carry-over items" so the work can move on. Only newly introduced bugs and regressions, security problems, and serious findings that are still unfixed are never relaxed, no matter the tier. Each loop also has an iteration limit (3 by default), so it always ends eventually. If a ticket still hasn't passed after reaching the limit, it stops as "blocked" and waits for morning.
The gatekeepers are all different elves, too. An elf that reads the code, an elf that hunts for gaps in the tests, an elf that eyes every input with an attacker's suspicion. Different elves look at the same deliverable at the same time. What one elf misses, the elf next to it picks up. The workshop is quiet at night, but on the workbench, a fairly strict inspection is going on.
Even back when I was developing with the normal, pre-Autopilot flow, most of what I did at the approval gates was rubber-stamping. I'd open a deliverable that had already made it through stage after stage of review, nod, and click Approve. I almost never sent anything back.
With this many gates stacked up, there isn't much left for a person to catch at the final approval. Human judgment has become nearly unnecessary. That is my honest impression after using Autopilot.
My job has shrunk to just this: looking at the PR in the morning and deciding whether to merge it. (You can also configure it to stop at creating a branch, or to go all the way and merge.)
The Elves File New Tickets
Using tree mode (/graph-ops:autopilot-tree) took the quality up another notch.
While the elves are working, they sometimes find problems that are separate from the main task: a performance issue spotted by a review, a gap noticed during testing. These are the kinds of things a person would jot down as "I'll do it later" and then promptly forget.
The elves don't let them slide. Once a ticket has been released, they gather up the reviews' carry-over items, the open issues from the report, and the follow-up notes from the implementation notes, and sort them one by one:
- Skip: Things outside the supported scope, harmless differences in wording, things that can't be verified without real hardware or an external environment, things that don't reproduce, and things already accepted in the plan or reviews. None of these become tickets.
- Child tickets: Things that should be dealt with within the scope of this ticket. They are filed as new tickets with the current ticket as their parent. The master elf picks those tickets up as part of the same run, and elves in other terminals take care of them.
- Backlog: Pre-existing problems that aren't part of the main task are bundled into a single backlog ticket and filed as a ticket with no parent. The master elf does not work on it during the run. It's a little something the elves leave behind.
Once an elf finishes a child ticket, it merges its branch into the parent's branch. Grandchildren's work goes into the child's branch, and the child's work goes into the parent's branch. Everything is merged from the bottom up, until it all comes together in the parent ticket's branch. The tickets in the tree are processed one at a time, in order.
"I'll do it later" actually gets done later. That's why tree mode improves quality.
A Record of One Day
One day, I decided to add a batch of new commands to a VS Code extension that collects commands for transforming selected text. It already had around 800 commands, so this was an extension of an existing feature.
That night, I filed the parent ticket and went to sleep with tree mode running. The ticket was to grow the 800 commands to 1,000.
The first elf took stock of the existing commands, designed which commands to add, checked them from a security standpoint as well, and then split the work into several tickets, filing them with its own ticket as the parent.
All through the night, terminals kept appearing on the screen, one after another. The master elf would pick up a ticket, hand the job to a lead elf, receive a short report, and reach for the next ticket. In a workshop with no one watching, that cycle went on all day long.
The elves implemented the twenty-odd commands listed in each ticket, and every time, code, QA, security, and non-functional reviews ran in parallel, followed by the documentation review and the approval gates.
The gates were doing their job. A process that took quadratic time. An expansion routine that could blow up explosively depending on the input. Tests that had been broken for a while. Everything the reviews found became grandchild tickets, and other elves fixed them. There were also problems caused by my local environment that weren't part of the main task, and those were left in the backlog as the elves' parting gift.
After working through a double-digit number of tickets, the run finished almost a full day later. The extension now had more than 1,000 commands, with nearly 200 newly added.
Bedtime Pays for the Time
Autopilot does not make the work faster.
Every gate it passes through triggers a review, and every rejection sends it around another loop. The more tickets are spun off, the more elves there are. It takes far longer than a person could sit there and watch.
But all of that happens at night, while the person is asleep.
By day, the shoemaker picks out the leather and decides what kind of shoes to make. In other words, he writes the ticket: what he wants to build, why he's building it, and what counts as done. As long as he writes that part carefully, the elves will work from that ticket through the night.
The hours we spend asleep were never producing anything to begin with. For work done in those hours, how long it takes hardly matters. In fact, precisely because you don't have to worry about time, you can add one more gate, or turn every crack found in the code and elsewhere into a ticket. You can go all in on quality without holding back.
Only one challenge remains: tokens.
The more gates and the more elves, the more tokens disappear. Just as the shoemaker had to pay for leather, the elves need to be fed. You'll have to decide how much to hand over based on your plan and your usage limits. This is the one cost that time can't cover for you.
Wrapping Up
When I wake up and open the terminal, an execution summary put together by the master elf is waiting for me. Which tickets got done? Which ones became child tickets, and which went to the backlog? Where did the elves approve something, and on what grounds? All of that can be read back in the Web UI. If any tickets stopped partway through (failed or blocked), that's where the shoemaker's work for the day begins.
On the workbench, the finished shoes are lined up. He picks them up and decides whether to put them in the shop. That's all the shoemaker's job is now.
Finally, let me sum up what I've covered:
- Thanks to layer upon layer of review gates, approval becomes mostly rubber-stamping, and human judgment is hardly needed anymore
- In tree mode, cracks found along the way get fixed as child tickets, and anything outside the main task is left in the backlog. "I'll do it later" actually gets done
- It takes time, but bedtime pays for it. That's why you can go all in on quality
- The only remaining challenge is token consumption
GraphOps is available on GitHub. Why not leave a ticket on your workbench tonight?
Top comments (0)