DEV Community

Cover image for How to Run an AI Agent Pilot That Actually Reaches Production
Pykero
Pykero

Posted on Originally published at pykero.com

How to Run an AI Agent Pilot That Actually Reaches Production

An AI agent pilot reaches production when it is scoped around one business metric, runs on your real data with a real stop date, and leaves behind receipts: run logs, an eval set, and a measured cost per completed task. Pilots that stall are almost never killed by the model. They stall because nobody wrote down what "done" meant before the demo looked impressive.

We have built agents for sales, support, document processing and internal operations, and we have watched the same pattern repeat. The pilot works in the room. Then three months pass and it is still "being evaluated". This guide is about how to avoid that as the buyer.

Why most pilots never leave the pilot stage

Three failure modes cover most of the cases we see, and each one has a specific shape:

  • The pilot measured the wrong thing. Teams measure "did the agent produce a plausible answer" instead of "did the ticket close, did the lead reply, did the clinician accept the draft". We made this mistake with our own cold-email engine: weeks spent grading how well the agent wrote the body copy, while the number that mattered, replies, sat flat. Plausible output is cheap. Outcomes are the point.
  • The pilot ran on clean data. Ten hand-picked documents or a curated list of questions. Production is the other ninety percent: scanned PDFs, Arabic mixed with English in the same message, customers who ask three things at once. A pilot that never sees those cases cannot tell you its real failure rate, and its "accuracy" number is a number about the demo set, not about your business.
  • Nobody owned the go/no-go decision. The vendor wants to extend, the champion is busy, and the pilot drifts. The tell is a pilot with a start date and no end date. That is not a pilot. It is a subscription, and the "three months of evaluation" we keep seeing is what it looks like from the inside.

Each of these is fixable at the scoping stage, which is where the buyer has the most leverage. The rest of this guide is the scoping checklist.

Step 1: Pick one metric and one workflow

Resist the urge to pilot "an AI assistant for the team". Pick a single workflow with a measurable outcome that already has a baseline. Examples that pilot well:

  • First-response drafts for one support queue, measured by agent acceptance rate and time to first response.
  • Lead qualification over WhatsApp for one product line, measured by qualified-lead handoff rate.
  • Invoice or claim data extraction for one document type, measured by field-level accuracy against the current manual process.

The metric must be something you would still care about if there were no AI involved. If you cannot compute the baseline today, spend the first week of the pilot computing it, because without a baseline the pilot can never prove anything. Our guide to calculating AI agent ROI covers how to turn those metrics into a number the board understands.

A lesson from our own tooling: in the cold-email engine we run for ourselves, we spent weeks tuning how the agent wrote the body of each email, and the metric barely moved. What moved it was changing the ask at the end. Replacing a calendar link with a soft "reply YES" lifted reply rates consistently, because scheduling a meeting from a cold email is more friction than an interested stranger will tolerate. The agent was fine. We had been measuring and optimising the wrong part of the funnel. Define the outcome metric first and let it tell you where to look.

Step 2: Run on real data, with real volume

Agree in writing how the pilot gets access to production-like data. This is usually the slowest part, so start the security and legal conversation on day one, not week three. Practical options, in order of preference:

  1. Shadow mode. The agent runs on live traffic but its output is logged, not sent. A human still does the job. You compare the two.
  2. Assisted mode. The agent drafts, a human approves or edits before anything goes out. Edit distance becomes a free quality signal.
  3. Scoped autonomous mode. The agent acts on its own for a narrow, low-risk slice, with clear escalation paths back to a person.

Shadow mode is underrated. It removes almost every objection from compliance, it generates the eval dataset for free, and it tells you the real failure rate before a single customer sees the agent. For healthcare and government buyers in the Gulf, it also sidesteps most of the data residency debate during the pilot, because nothing leaves the existing workflow.

Step 3: Define the receipts before you start

Every "done" in a pilot needs a receipt, something you can inspect after the fact. Before the build starts, agree on exactly what artefacts the pilot produces:

  • Run logs for every task. Inputs, tool calls, model outputs, latency, token usage. Not a dashboard screenshot. Structured logs you can query. OpenTelemetry is the obvious neutral format if you want them to outlive the vendor relationship.
  • An eval dataset. A few hundred real cases with the expected outcome, scored automatically. This is the asset that lets you upgrade models, switch providers, or change prompts later without guessing. If the pilot ends without one, you have bought a demo.
  • A cost per completed task. Model spend, tool calls, retries and human review time, divided by tasks actually completed. Not cost per API call.
  • A failure taxonomy. The ten most common ways it went wrong, with counts. This becomes the production backlog.

Put these in the statement of work as deliverables. We cover how to make them contractual in putting evals in AI vendor contracts. A vendor who pushes back on delivering logs and eval data is telling you something.

Step 4: Set the go/no-go criteria and date

Write down three numbers before the first line of code:

  • Ship threshold. The metric at which you commit to production, for example "agent drafts accepted without edits on 60 percent of tickets, zero policy violations in 500 runs".
  • Kill threshold. The level below which you stop, regardless of how promising the roadmap sounds.
  • Decision date. A calendar day on which the sponsor decides, with the data in front of them.

The space between ship and kill is where most pilots live, and it is where you need a rule. Ours is simple: one extension of fixed length, with a specific hypothesis about what will move the metric. If the second run does not clear the ship threshold, stop. Open-ended extensions are how a pilot quietly becomes a year of spend.

Step 5: Price the pilot so it is worth finishing

A free pilot has no deadline, because the vendor's incentive is to convert you, not to finish. A pilot priced at the full production rate makes the buyer nervous about committing. The structure that works in our experience is a fixed-price pilot that explicitly includes the eval harness and logging, followed by a separately scoped production phase. The buyer owns everything the pilot produces, whichever way the decision goes. The trade-offs between fixed and variable pricing are covered in fixed price vs time and materials.

Make sure the pilot also answers the architecture questions that affect production cost. In our own outreach engine, a pilot-scale comparison showed that a single model call which extracted the facts from a scraped page and drafted the email in one step beat a multi-step chain on both cost and output quality. Had we only tested the chain, we would have carried its cost into production. Use the pilot to run at least one such A/B on the core design, while the volume is small enough to make it cheap.

What production hardening actually adds

Buyers often ask why the production build costs multiples of the pilot when "it already works". The honest list, and where each item comes from if the pilot left proper receipts:

  • Retries, idempotency and dead-letter handling for every tool the agent calls. The failure taxonomy tells you which tools actually fail and how often, so you harden the top of the list first instead of wrapping everything.
  • Rate limits and budget caps so a loop cannot run up a bill overnight. The cost per completed task from the pilot is the number the cap is set against. Without it, the cap is a guess.
  • Permission scoping, so the agent only sees what the current user may see. In shadow mode the agent saw everything the human saw. Production has to narrow that, and the run logs show which data it actually used.
  • Monitoring for silent failures, not just errors. An agent that stops running still returns a 200. The eval set, replayed on a schedule, is the cheapest alarm for a model or prompt that has quietly degraded.
  • Human review queues and audit trails, which regulated buyers will need before sign-off. If the pilot ran in assisted mode, the edit-distance data already tells you how large that queue needs to be and which cases should land in it.

None of this is visible in a demo, and all of it is what separates a pilot that worked from a system you can put in front of customers. A pilot that produced real logs, a real eval set and a real failure taxonomy makes this list shorter and far more predictable to quote. A pilot that produced only a demo means the production team starts by rebuilding the pilot to find out what it should have measured.

A one-page pilot charter

If you take one thing from this article, write a single page before the kickoff with: the workflow, the outcome metric and its baseline, the data access mode, the four receipts, the ship and kill thresholds, the decision date, and who owns the decision. Share it with the vendor and the sponsor. A pilot run against that page either ships or teaches you something specific. Either result is worth paying for.

If you are planning an AI agent pilot and want a partner who will put those thresholds in the contract rather than argue about them later, let's talk.


Originally published on the Pykero blog.

Top comments (0)