DEV Community

Ram Ji Tripathi
Ram Ji Tripathi

Posted on Originally published at blog.ramjwork.in

Designing Human-in-the-Loop AI: Why an AI-Generated Reply Should Be a Draft, Not an Action

1. The Dangerous Shortcut Is Treating Generation as Action

An LLM returning plausible text is a different event from an application committing that text to a customer conversation. The first produces a candidate. The second changes application state under an authenticated user's authority. Connecting the two is an architectural decision, not an inevitable next step after generation.

It is tempting to connect a successful generation response directly to a message-creation request. The integration is short: get text, submit text, display success. But that connection quietly turns a candidate answer into an application decision. It also makes provider completion the trigger for a user-visible side effect.

In my AI Support Assistant, generation returns text to a suggestion preview. A person chooses Use Reply to move it into an editable composer, then separately chooses Send to submit the current content through the message API. The model supplies a candidate; the application keeps commitment behind a different operation.

This is a specific choice for a support-reply workflow. It does not establish that every AI operation needs permanent manual approval. It gives me a concrete place to examine authority: who can propose a change, who controls the working copy, and which request commits it?

2. The Boundary in My AI Support Assistant

The relevant screen is a ticket detail page with a conversation, an AI Reply Assistant, and a message composer. React handles the interface; an Express backend generates suggestions and persists messages. That is enough application context for this discussion. The broader system is covered in How I Architected an AI Customer Support System.

The exact interaction matters. Generate AI Reply requests a suggestion and displays it in a separate preview. It does not automatically fill the composer. Use Reply transfers the displayed text into the composer. Regenerate requests another candidate. Send submits the current composer content.

A user can leave a suggestion unused, regenerate it, or use and edit it. The main workflow controls are Generate AI Reply, Use Reply, Regenerate, and Send. Use Reply changes client working state; it does not approve a message for immediate sending.

The backend generation controller reads messages for the authorized ticket in ascending creation order. It formats each as ROLE: content and passes that conversation to the reply service. Although the request identifies a ticket, this path does not add the ticket title or description to the prompt. An unsent composer draft is also absent: the generation request sends a ticket identifier, not the current input text.

3. Three Different States: Output, Draft, and Action

The reusable model is model output → human-owned working draft → authoritative side effect. Use Reply crosses the first boundary. A successful message creation crosses the second; clicking Send attempts that operation. The words may initially be identical, but their ownership and consequences change.

Stage Representation in this implementation Owner / source of truth Consequence
Model output Returned suggestion, displayed in the AI panel Generation response held by the mutation Candidate text; no conversation message created
Human-owned working draft Composer content after Use Reply Local React state User can change or erase the working copy
Authoritative side effect Message persisted through the message API Backend and persisted message record Successful creation adds a saved message

“Human-owned” describes control of the draft, not a claim that the person authored every word. The user might send the suggestion unchanged. The important distinction is that edits and commitment belong to the application workflow rather than being implicit consequences of model output.

TanStack Query holds the generation mutation result and the message query cache. That does not make both values equivalent server-domain records. A suggestion response is not a saved message, and a cached message is a client representation of a backend record.

The composer uses local state initialized from the selected suggestion. This ownership distinction overlaps with React State Architecture, but the concern here is authority. Choosing local state is useful because unfinished text should remain working input until a separate operation commits it.

4. Generation and Sending Are Separate Operations

The frontend generation function posts to an AI-specific endpoint. This function excerpt omits imports and the timeout constant, which is set to 60 seconds:

export async function getReplySuggestion(ticketId: string) {
  const response = await api.post<ApiResponse<ReplySuggestionResult>>(
    `/ai/tickets/${ticketId}/reply-suggestion`,
    undefined,
    { timeout: REPLY_SUGGESTION_TIMEOUT_MS },
  );
  return response.data.data;
}
Enter fullscreen mode Exit fullscreen mode

Source: frontend/src/features/ai/api/ai.api.ts, getReplySuggestion.

The corresponding hook is a generation mutation. The reply service returns suggestion and tokensUsed; it does not call message creation. OpenAI is the primary provider, with an optional configured Gemini fallback for eligible failures. Either provider still supplies candidate text to the same workflow.

The generation controller's sequence is visible in this simplified excerpt. Parameter extraction, the message query details, and response wrapping are omitted:

await getAuthorizedTicketService(ticketId, {
  userId: req.user!.userId,
  role: req.user!.role,
});
// Read ticket messages, ordered by createdAt ascending.
const conversation = formatConversation(messages);
const aiResult = await generateReplySuggestion(conversation, {
  ticketId,
  requestId: typeof req.id === "string" ? req.id : undefined,
});
Enter fullscreen mode Exit fullscreen mode

Source: backend/src/modules/ai/ai.controller.ts, generateReplySuggestionController.

Sending uses a different frontend function and mutation, posting content to /tickets/${ticketId}/messages. Its backend controller derives the sender from the authenticated request and invokes the message service. These endpoint paths are relative to the application's API base.

Generation calls an external provider, produces operational logs, and performs a ticket-access lookup that can populate a cache. The boundary concerns a specific side effect: this suggestion path does not create or send a customer message. Provider activity and operational writes still happen, but they do not add the candidate to the conversation.

For a future extension, I would preserve this separation in the service contract as well as the screen. A caller asking for a suggestion should receive candidate data, without needing to know how conversation writes work. A caller committing a message should supply the intended content and pass the application’s authorization checks. Moving the same feature to another client should not accidentally change the meaning of generation into permission to send. This is a design recommendation derived from the current boundary, rather than a claim about additional clients in my application.

Architecturally, separate operations let completion mean different things. Generation completion means a candidate is available. Message completion means the application has accepted a send request and returned a created record. Neither status should silently stand in for the other.

5. Human Review Is Part of the Architecture, Not Just the UI

The handoff is explicit in the page component. These are two excerpts from the same component, with layout omitted:

onUseReply={(content) => {
  setComposerDraft(content);
  setComposerKey((current) => current + 1);
}}

<MessageComposer
  key={composerKey}
  ticketId={data.id}
  initialContent={composerDraft}
  onMessageSent={() =>
    setMessageRefreshToken((current) => current + 1)
  }
/>
Enter fullscreen mode Exit fullscreen mode

Source: frontend/src/features/tickets/pages/TicketDetailPage.tsx, TicketDetailPage.

Changing the key remounts the composer, whose local content starts from initialContent. If a person has already typed a draft, clicking Use Reply replaces that content. The current handoff offers neither a merge nor a confirmation dialog. That is a UX limitation of adopting a suggestion, even though the send boundary remains separate.

Regenerating alone does not invoke this callback, so a new preview does not automatically replace the composer. The replacement happens when the user chooses Use Reply. That detail matters when explaining what control the person actually has.

Inside the composer, the textarea's value comes from local content, and its change handler updates that state. The send handler reads the current value, so user edits become the submitted text. There is no requirement that it continue matching the original suggestion.

The sequence gives a person two distinct decisions: whether to adopt the candidate as working text, and whether to commit the current draft. They can send it unchanged or click through without carefully reading. The server receives content and an authenticated sender, without proof of deliberation. Human control is explicit; meaningful review is still a behavior the interface cannot prove.

6. Follow One Reply Through the System

Consider an authorized support agent looking at a ticket assigned to them. They request a reply. The backend checks ticket access, reads persisted conversation messages, and generates candidate text. The frontend displays that result without creating a conversation entry.

Authorized generation request
        ↓
Persisted ticket-message context
        ↓
AI generation → returned candidate
        ↓
Suggestion preview
        ↓
Explicit Use Reply
        ↓
Editable composer (local working draft)
        ↓
Explicit Send (separate message mutation)
        ↓
Message API: authentication + ticket authorization
        ↓
Persisted message (authoritative conversation state)
        ↓
Ticket timestamp update + summary enqueue
        ↓
Successful response → append record to query cache
Enter fullscreen mode Exit fullscreen mode

This is the current reply flow, including the service work between persistence and response. It is a workflow diagram, not a claim that every arrow is atomic.

The composer validates nonblank input and submits the trimmed current content. This simplified excerpt preserves the success and failure behavior:

try {
  await createMessageMutation.mutateAsync({
    content: content.trim(),
  });
  setContent("");
  setFieldError(undefined);
  onMessageSent?.();
} catch (error) {
  setSubmitError(
    getTicketErrorMessage(error, "Unable to send message. Please try again."),
  );
}
Enter fullscreen mode Exit fullscreen mode

Source: frontend/src/features/messages/components/MessageComposer.tsx, handleSubmit; validation precedes this block.

After a successful response, the message hook appends the returned record to the ticket's query cache:

onSuccess: (newMessage) => {
  queryClient.setQueryData<Message[]>(
    messageKeys.ticket(ticketId),
    (existing) => (existing ? [...existing, newMessage] : [newMessage]),
  );
},
Enter fullscreen mode Exit fullscreen mode

Source: frontend/src/features/messages/hooks/useMessages.ts, useCreateMessage.

The cache receives the returned record after success. There is no optimistic message insertion in this hook. The page also increments a refresh token, which the thread uses for scrolling rather than automatic query refetching. This reconciles the returned message into the local list; it does not establish complete consistency across clients.

7. Failure Becomes Easier to Contain When Suggestion Is Not Action

If generation fails, the AI panel displays an error and offers Try Again. Its mutation error does not call the composer handoff or the send mutation. The generation error path leaves the existing composer content untouched. The interface also handles a returned suggestion that becomes empty after trimming, offering another attempt.

If a candidate is never used or sent, it does not become a conversation message through this path. The suggestion panel resets its mutation state when the ticket identifier changes. The inspected components do not implement durable draft saving. An unused suggestion or unsent draft is not a persisted review record in this workflow.

If sending fails, the composer catches the error and displays a send failure. Its content is cleared only after the awaited mutation succeeds, so the draft remains available on the error path. While sending is pending, the textarea is disabled. These are useful local behaviors, but they do not settle what happened on the server.

The message service first persists the message, then updates the ticket timestamp, then awaits a queue enqueue for summary work. These steps are not wrapped in a transaction in this function. If the timestamp update or enqueue throws, the request can fail after the message has been saved. The client retains its draft, but that does not mean the conversation is unchanged.

More generally, a lost response can also leave a client uncertain about a completed write. The practical distinction is between failure to receive confirmation and confirmed absence of a side effect. A client-visible send error does not necessarily prove that message creation did not happen.

A stronger send design could add idempotency and deliberate recovery for uncertain outcomes. That is proposed work, not an existing guarantee. The important current separation is that retrying generation does not itself create another customer message. Retrying a send belongs to a different failure domain and needs different reasoning.

Keeping suggestion separate from action contains a generation problem at the candidate stage. Send still has its own persistence and recovery concerns. This separation gives each failure a clearer meaning, while leaving uncertain writes to be handled explicitly in the message workflow.

8. Authorization Still Belongs on the Server

The ticket page shows the AI panel to agents and administrators. That is a presentation rule. A browser user can construct requests independently of which controls are visible, so the backend must decide what each request is allowed to do.

The AI route authenticates the request, applies its rate limiter, and restricts generation to AGENT and ADMIN roles. The controller then checks access to the particular ticket. The ticket authorization policy allows administrators access across tickets, agents access to assigned tickets, and customers access to their own tickets. Customers still cannot generate replies because the AI route has the additional role restriction.

Message creation has different permissions. Its route authenticates and validates content, while its service checks ticket access before persistence. It does not have the AI route's staff-only role gate. Customers can create messages on their own tickets; agents on assigned tickets; administrators across tickets.

The essential check is short:

export const createMessageService = async (data: CreateMessageInput) => {
  await getAuthorizedTicketService(data.ticketId, data.accessContext);
  const message = await createMessageRepository({
    content: data.content,
    ticketId: data.ticketId,
    senderId: data.senderId,
  });
  // Ticket timestamp update and summary enqueue follow before return.
};
Enter fullscreen mode Exit fullscreen mode

Source: backend/src/modules/messages/message.service.ts; simplified excerpt, not the complete function.

The controller takes senderId from authenticated identity rather than a body-supplied author. The shared ticket service checks access on both its cached-ticket and database lookup paths. These checks establish application authority for the request; they do not validate the factual truth of the message.

Separating generation permission from message permission also keeps the capabilities understandable. Being allowed to obtain a candidate and being allowed to add a message are distinct decisions, even when the same staff member can do both.

9. What This Design Does Not Solve

The reply prompt asks for concise, polite, solution-oriented text and instructs the model not to invent policies, promise refunds, or fabricate technical actions. Those are prompt instructions. The inspected reply path does not independently verify the answer against an authoritative policy source or enforce those instructions with a deterministic validator.

There is no reply confidence threshold or factual-verification stage in this path. Human editing can correct an answer, but human review does not guarantee correctness or eliminate hallucinations. The implementation also does not establish that every suggestion is reviewed before submission.

Context can become stale. Generation reads persisted messages at request time; the user may send later, after the conversation has changed. The inspected send request contains content, without a conversation-version precondition binding it to the generation snapshot. An empty message history also means this generation path has no conversation text; identifying the ticket does not automatically supply its description.

Provenance is a separate concern. AI operation logs include operational context and outcomes, but the inspected reply path does not establish a persisted linkage showing that a sent message originated from a particular AI suggestion. The Message model has no suggestion reference or AI-authorship field. Existing AI interaction storage elsewhere does not establish that linkage for this workflow.

The human boundary does not establish policy compliance, production safety, complete observability or provenance, or measured reliability improvements. The inspected reply path also has no automated policy engine or escalation step. What this implementation demonstrates is narrower and useful: a person has control over the working text before a separately authorized request attempts to commit it.

10. When Would I Allow More Automation?

The following framework is proposed guidance for future designs. It is not implemented in this application. I would evaluate the action being committed, rather than treating every model output as equally risky.

Generic action Reversible? Potential harm Validation available? Human approval approach
Save a private draft Usually easy to revise Lost or misleading working text Schema and storage checks May be unnecessary if replacement is controlled
Suggest an internal category Usually reversible Misrouting or delayed handling Allowed categories and routing rules Depends on impact and recovery
Send a customer reply Cannot retract what was read Incorrect advice or commitments Approved facts and policy checks may help Retain review until evidence supports narrower automation
Issue a refund Sometimes difficult to reverse Financial loss or inconsistent treatment Account, amount, and eligibility rules Require explicit authority appropriate to the risk
Delete customer data Often hard to reverse Data loss and account harm Identity, scope, retention rules Strong approval and recovery requirements

Reversibility is only one dimension. A reversible change applied to thousands of accounts may have a larger blast radius than one irreversible but tightly bounded action. I would examine scope, domain risk, and the cost of a wrong action together.

Validation should come from something stronger than the model's confident wording. Deterministic checks can constrain amounts, destinations, allowed operations, and required permissions. Trusted domain data can support specific factual checks. Neither approach automatically verifies every sentence, but each can reduce uncertainty about a defined commitment.

I would also require observability and a recovery plan: which input led to the action, who or what authorized it, how failures are detected, and how operators can intervene. If human review is unavailable, that does not itself justify removing the gate; it changes the service and fallback design.

Before widening automation, I would define what evidence could overturn the decision to require review. That might include evaluation against representative cases, checks for forbidden commitments, and a demonstrated way to detect and recover from mistakes. I would also define conditions that return the workflow to human handling, such as missing domain information or an action outside the permitted scope. These are future decision criteria; the current application does not implement this automation framework.

Automation can then expand within a bounded class of actions, with evidence appropriate to that class. A decision to automate private drafts should not silently authorize refunds or customer-facing messages. Permission should follow the side effect, with separate evaluation when its consequence changes.

11. A Practical Human-in-the-Loop Design Checklist

For an AI-assisted workflow, I would ask these questions before connecting generation to commitment:

  • What exactly does the model return, and which operation creates the domain record?
  • Where can a person inspect and edit the candidate before commitment?
  • Does adopting a suggestion replace existing work, and is that behavior clear?
  • Which server checks authorize generation and which authorize the action?
  • What context was used, and can it change before commitment?
  • What does each error prove, particularly after partial persistence?
  • Does the client reconcile a confirmed record or merely assume success?
  • What evidence would be needed to automate this particular side effect?

These questions are useful even when the model and prompt change. They expose responsibilities that belong to the application: working-state ownership, permissions, persistence, recovery, and the meaning of a successful response.

For my current reply workflow, the answer is an editable draft and a separate send request. Strengthening draft replacement, uncertain-send recovery, and provenance would be additional work. The existing boundary gives those improvements a clear place to attach without pretending they are already present.

12. Closing

In my reply workflow, Use Reply makes candidate text editable working state. Send separately asks the backend to create the message under the authenticated user's authority. The model's output supplies content; the application defines the commitment boundary.

When integrating AI, ask who has authority to commit the side effect, then make that authority explicit in the interface and the server operation.


Originally published on my engineering blog: https://blog.ramjwork.in/ai-engineering/human-in-the-loop-ai-reply-design

Top comments (2)

Collapse
 
anciwasim profile image
Wasim Sheikh •

Splitting Use Reply from Send is the right seam. The part I'd poke at is the Send itself, since that's now the only real side effect. If the message API times out, does the client retry? We had agents (and people double-clicking) post the same reply twice because a timeout on a write got treated as "failed" when it had actually landed. An idempotency key generated when the draft enters the composer fixed it for us, and a timeout became "unknown, go check" instead of "try again". Is that something you handle on the Express side yet?

Collapse
 
ramji_tripathi_095c7f4810 profile image
Ram Ji Tripathi •

Great point. That’s exactly the uncomfortable edge case in the current design. The Express-side send path doesn’t currently have an idempotency key tied to the draft/send attempt, so I wouldn’t want the client to blindly treat a timeout as “safe to retry.”
The next step I’d take is very similar to what you described: generate an idempotency key for the send attempt, persist/deduplicate against it on the server, and let an ambiguous timeout become “check the outcome” rather than “send again.”
I deliberately kept that separate from the human-in-the-loop boundary in this article, but it’s an important reliability problem once the human actually commits the side effect. Thanks for calling it out