DEV Community

Mary Khachatryan
Mary Khachatryan

Posted on Originally published at wagglet.com AI-assisted

Stop Calling an Agent Run “Done”: Model Delivery, Acceptance, Merge, and Deployment Separately

An agent says the task is finished. A branch exists. CI is green. The ticket moves to Done.

Those four facts often arrive close together, but they are not the same fact. Treating them as one creates an avoidable ambiguity: did the agent finish an attempt, did a person accept the result, did the code land, or did users receive it?

A safer workflow gives each question its own state transition, permission check, and receipt.

This article builds that workflow as a small state machine. The example is intentionally implementation-level: named actors, guarded transitions, idempotent delivery, and repository evidence that cannot silently overwrite review state.

Start with two kinds of truth

The work tracker owns the truth about intent and review. The repository and release system own the truth about integration and deployment.

Do not force them into one status field.

type WorkState =
  | "draft"
  | "open"
  | "claimed"
  | "delivered"
  | "accepted";

type IntegrationState =
  | "not_applicable"
  | "unreported"
  | "branch_reported"
  | "pull_request_open"
  | "merged"
  | "blocked";

type DeploymentState =
  | "not_applicable"
  | "not_deployed"
  | "deploying"
  | "deployed"
  | "failed";
Enter fullscreen mode Exit fullscreen mode

An accepted task can still be unmerged. A merged change can still be undeployed. A non-code task can be accepted without either field applying.

That is not inconsistency. It is accurate modeling.

Define the actors before the transitions

A lifecycle is only useful if the system knows who may move it.

type Role =
  | "requester"
  | "task_owner"
  | "runner"
  | "reviewer"
  | "repository_maintainer"
  | "release_owner";

type Actor = {
  id: string;
  roles: Role[];
};
Enter fullscreen mode Exit fullscreen mode

One person may hold several roles, but the permission check should still name the role being exercised. That keeps an administrator override from looking like an ordinary peer review.

For higher-risk work, require the reviewer to be different from the runner. If self-verification is allowed, record it explicitly rather than presenting it as independent acceptance.

Make the happy path boring

The main workflow can stay small:

draft -> open -> claimed -> delivered -> accepted
Enter fullscreen mode Exit fullscreen mode

Each transition answers one question:

Transition Question answered Required evidence
draft → open Is the task safe and specific enough to offer? Prepared brief and owner
open → claimed Who owns this attempt? Eligible runner and attempt ID
claimed → delivered What result did the runner produce? Delivery report and artifacts
delivered → accepted Does the result satisfy the task? Reviewer verdict

The state machine should reject shortcuts. A runner cannot accept their own delivery under the ordinary review path. A repository webhook cannot move a work item to Accepted. A successful deployment cannot invent a missing delivery report.

Guard every transition

The transition function should validate both the current state and the actor.

type Ticket = {
  id: string;
  state: WorkState;
  ownerId: string;
  runnerId?: string;
  activeAttemptId?: string;
  revision: number;
};

function assertCanDeliver(ticket: Ticket, actor: Actor) {
  if (ticket.state !== "claimed") {
    throw new Error("Only claimed work can be delivered");
  }
  if (ticket.runnerId !== actor.id) {
    throw new Error("Only the active runner can deliver this attempt");
  }
  if (!ticket.activeAttemptId) {
    throw new Error("Claimed work must have an active attempt");
  }
}

function assertCanAccept(ticket: Ticket, actor: Actor) {
  if (ticket.state !== "delivered") {
    throw new Error("Only delivered work can be accepted");
  }
  if (!actor.roles.includes("reviewer")) {
    throw new Error("Acceptance requires reviewer permission");
  }
  if (ticket.runnerId === actor.id) {
    throw new Error("Independent acceptance requires another person");
  }
}
Enter fullscreen mode Exit fullscreen mode

These checks belong on the server. Hiding a button in the interface is useful feedback, not authorization.

Treat delivery as a report, not a verdict

Delivery means “this attempt is ready for judgment.” It does not mean the work is correct.

A delivery record should be append-only and tied to the active attempt:

type Delivery = {
  id: string;
  ticketId: string;
  attemptId: string;
  runnerId: string;
  summary: string;
  branch?: string;
  commitSha?: string;
  createdAt: string;
  operationId: string;
};
Enter fullscreen mode Exit fullscreen mode

operationId makes retries idempotent. If a client times out after the server stores the delivery, sending the same operation again should return the original result instead of creating a second delivery.

Use an expected ticket revision as well:

async function deliver(input: {
  ticketId: string;
  expectedRevision: number;
  operationId: string;
  summary: string;
  branch?: string;
}) {
  // 1. Return the existing result for this operationId.
  // 2. Load the ticket and compare its revision.
  // 3. Check state and runner permission.
  // 4. Insert the delivery and transition atomically.
  // 5. Return the new revision.
}
Enter fullscreen mode Exit fullscreen mode

This prevents a stale browser tab or delayed agent process from delivering against an attempt that has already been reopened or reassigned.

Support two different rework paths

Review failure is not one situation.

If the same runner should fix the result, send it back to Claimed and keep ownership. Start a new attempt or revision while preserving the previous delivery as history.

If another runner should take over, reopen it to Open and clear the claim. The previous delivery remains evidence, but it is no longer the standing answer.

type ReviewDecision =
  | { kind: "accept"; note?: string }
  | { kind: "send_back"; feedback: string }
  | { kind: "reopen"; reason: string };
Enter fullscreen mode Exit fullscreen mode

Do not implement both choices as “move the card left.” They differ in ownership, notifications, attribution, and what context the next run needs.

Keep repository updates on a separate rail

When a delivery reports a branch or commit, store that as a claim first. Then verify it against the repository provider.

type RepositoryEvidence = {
  reportedBranch?: string;
  verifiedHeadSha?: string;
  pullRequestUrl?: string;
  mergeState: IntegrationState;
  checkedAt?: string;
};
Enter fullscreen mode Exit fullscreen mode

A repository callback can update mergeState, but it must not set state = "accepted". Likewise, accepting the work must not set mergeState = "merged" unless the repository confirms it.

This separation matters when:

  • branch protection requires another review;
  • CI fails after the task review;
  • the base branch advances and creates a conflict;
  • a pull request is intentionally left open for batching;
  • a deployment rolls back after a successful merge.

The UI can present these facts together without collapsing them. For example:

Work: Accepted
Repository: Pull request open
Deployment: Not deployed
Enter fullscreen mode Exit fullscreen mode

That line tells a release owner exactly what remains.

Record events, not just the latest values

Current state is useful for the board. Events explain how the work arrived there.

type TicketEvent = {
  id: string;
  ticketId: string;
  actorId: string;
  type:
    | "published"
    | "claimed"
    | "delivered"
    | "sent_back"
    | "reopened"
    | "accepted"
    | "merge_state_changed"
    | "deployment_state_changed";
  attemptId?: string;
  metadata: Record<string, unknown>;
  createdAt: string;
};
Enter fullscreen mode Exit fullscreen mode

An audit log makes several important distinctions visible: a failed attempt versus a bad task description, self-verification versus independent review, and an accepted result versus a completed release.

It also gives teams data they can improve without inventing a single “agent success” score.

A practical implementation checklist

Before shipping this lifecycle, verify that:

  1. Every mutation checks the current state on the server.
  2. Claim and delivery are attached to an attempt ID.
  3. Mutations use an idempotency key and expected revision.
  4. Delivery stores evidence but does not imply acceptance.
  5. Send-back and reopen have different ownership behavior.
  6. Acceptance records who reviewed the work and whether review was independent.
  7. Branch, pull request, merge, and deployment states come from their actual systems.
  8. Repository and deployment updates cannot overwrite the work-review state.
  9. The event log preserves superseded deliveries and rework reasons.
  10. The interface shows the separate states together in plain language.

If you want to compare this model with a working product lifecycle, Wagglet documents a request-to-delivery implementation for project management software for AI agents. The useful idea is not a particular board layout. It is refusing to let one convenient “Done” label stand in for four different decisions.

When delivery, acceptance, merge, and deployment remain separate, automation becomes safer and human review becomes easier to audit. The workflow gains a few explicit transitions, but the team loses a much larger amount of guesswork.


Disclosure: I’m publishing this on behalf of the Wagglet team. AI assistance was used to structure and edit this article; I reviewed the technical content and final wording.

Top comments (0)