DEV Community

Cover image for The Month-Two Checkpoint: When to Pull the Plug on an AI Pilot
Sonal Jain
Sonal Jain

Posted on

The Month-Two Checkpoint: When to Pull the Plug on an AI Pilot

Every AI pilot I run has a checkpoint around week eight where we decide, in writing, whether to continue, change direction, or stop. The decision itself is rarely hard. What makes it possible is that the criteria were written down in week one, before anyone had a favourite outcome to defend.

Why week eight?

Week eight is late enough for real users to have hit the system with real data, and early enough that stopping is still cheap. By then the demo glow has faded, the first bill has arrived, and the intended users have either adopted the tool or found reasons not to.

Sunk cost has not hardened into pride yet. Wait until month five and the pilot has a budget line, a name in a board deck and a team whose next quarter depends on it. Nobody stops those. They just get quietly extended.

What are the criteria, and who writes them?

The criteria are written by the delivery lead and the sponsor together at kickoff, on one page, in four groups: the outcome and the number that counts as working, adoption without chasing, cost per unit and its direction, and how many of the starting unknowns are still unknown.

Criterion The question it asks Example stop condition
Outcome What number counts as working? Fewer than a third of drafts go out unedited by week eight
Adoption Are people using it without being chased? Weekly active users flat or falling for three weeks
Unit cost What does one unit cost, and which way is it moving? Cost per handled ticket still above the manual baseline
Unknowns How many week-one unknowns are still unknown? The biggest question at kickoff is still unanswered

A support drafting pilot might set its outcome line at "the agent's draft goes out with no edits at least half the time". Putting the stop conditions next to the success conditions is the whole trick. In week eight nobody argues about them, because they were agreed by the same people in week one, before there was anything to be attached to.

What does the conversation sound like when it isn't working?

Direct and early. The checkpoint meeting is booked on day one, so it exists before anyone is nervous about it, and it opens with the same one-page criteria. The recommendation is stated plainly, with the measured number, the agreed floor, and the direction of travel.

The running order I use:

  1. Read the week-one page out loud, success conditions and stop conditions together.
  2. Put the measured numbers next to them, with the trend rather than the latest value.
  3. State a recommendation, whether that is continue, narrow or stop, before opening the floor.
  4. Write the decision and its date on the same page, and name who owns the next one.

In practice step three sounds like this: "We are at 28 percent unedited, we agreed a third was the floor, and the trend is flat. My recommendation is to narrow it to refund queries only, where it is at 61 percent, and review again in four weeks."

I take seriously the industry warnings that a large share of AI agent projects will be dead by 2027, and my reading of that number is that most of those projects should have died in month two, cheaply and with everyone's dignity intact, rather than in month fourteen with a budget attached. Killing a pilot at the checkpoint is a good outcome. It is the reason the checkpoint exists.

What does a pivot usually look like?

Smaller, almost always. A pivot after a failed checkpoint narrows the scope to the slice where the numbers were already good, drops the autonomy level so the tool suggests instead of acts, adds a human checkpoint in front of it, or changes the data the pilot runs on.

That last one comes up more than people expect, because the pilot ran on the data that was easy to get rather than the data the job needed. Whichever it is, the pivot gets its own week-one page and its own checkpoint. A pilot that pivots without new criteria is a pilot that has learned to avoid being measured.

The first hour of any AI consulting engagement I run now goes on stop conditions rather than on the roadmap, because the roadmap is the easy part and nobody disagrees with it yet.

At Shanti Infosoft I would rather run four small pilots that each end with a clear decision than one long pilot that ends with a shrug. The decisions are the deliverable, even the ones that say stop. If a pilot of yours is overdue a checkpoint, my calendar is here.

If your current AI pilot were being judged next week, what would it be judged against, and who agreed to that?

Sonal Jain runs delivery at Shanti Infosoft, a CMMI Level 5 team that has built software for more than 700 companies.

Top comments (0)