DEV Community

Manoranjan Rajguru
Manoranjan Rajguru

Posted on

Human-in-the-Loop Approvals in Microsoft Foundry: Pausing Long-Running Agents Indefinitely Without Losing State

Human-in-the-Loop Approvals in Microsoft Foundry: Pausing Long-Running Agents Indefinitely Without Losing State

Day 16 of Foundry 100 Days / 100 Blogs

Table of Contents

  1. The Problem: Agents That Need a Human, Not a Retry
  2. Why This Matters
  3. Core Concepts: Task IDs, Entry Modes, and Suspension
  4. Architecture: Where the Approval State Actually Lives
  5. Implementing an Approval Turn Step by Step
  6. Wiring the Human Side: Notifications and Resume Triggers
  7. Framework Interrupts: LangGraph and Microsoft Agent Framework
  8. A Real-World Scenario: Expense Approval With Escalation
  9. Production Considerations
  10. Security Considerations
  11. Performance and Scale
  12. Cost Considerations
  13. Common Mistakes and Pitfalls
  14. Alternatives and Trade-offs
  15. Practical Recommendations
  16. Conclusion
  17. References

The Problem: Agents That Need a Human, Not a Retry

Most agent failure modes are things you engineer around: a tool call times out, a model hallucinates a malformed JSON payload, a rate limit trips. You retry, you fence the side effect, you fall back to a smaller model. Microsoft Foundry's resilient task subsystem — leases, checkpoints, entry_mode recovery — exists precisely to make those failures invisible to the user (see Day 1 of this series, on crash-resilient long-running agents).

But there's a category of "interruption" that isn't a failure at all: the agent is working correctly, and it has simply reached a point where it is not allowed to keep going without a person saying yes. A finance agent that wants to submit a $1,200 expense reimbursement. A DevOps agent that wants to run terraform apply against production. A support agent that wants to issue a refund above a threshold. In every one of these cases, the "right" behavior is not to fail, retry, or guess — it's to stop, ask, and wait, for however long it takes a human to look at a Slack message, a ticket queue, or an approval inbox. That could be ninety seconds. It could be three days if the approver is on vacation.

Most agent runtimes handle this badly. If your agent is a synchronous request/response call sitting behind an HTTP connection, you cannot hold that connection open for three days. If you fake it with polling and an external state machine (a Durable Function, a Temporal workflow, a hand-rolled Postgres table with a status column), you've now built and are maintaining a second orchestration layer outside the agent runtime, with its own failure modes, and the agent's own conversational state (tool call history, prior turns, checkpoints) lives somewhere else entirely from the approval state. Reconciling the two after a crash is exactly the kind of glue code nobody wants to own.

Microsoft Foundry's Agent Service takes a different position: human-in-the-loop approval isn't bolted onto the long-running agent primitives as a separate feature — it's a natural consequence of how multi_turn_task chains already work. If you understood Day 1's resilience model, you already have 80% of the mental model for how approvals work. This article is the other 20%: how to actually build the pause point, how to drive it from an application, how it survives a crash mid-pause, and where the sharp edges are in production.

Why This Matters

Enterprise AI adoption keeps running into the same wall: organizations are comfortable letting an agent draft an action but not execute it unattended, especially for anything touching money, infrastructure, or customer-facing communication. Every serious agent framework has converged on some notion of "approval gate" — LangGraph has interrupt(), Microsoft Agent Framework has RequestInfoEvent and ApprovalRequiredAIFunction (which we touched on in Day 7's workflow migration piece), CrewAI has human input tools. What's different about Foundry's approach is that it doesn't treat the approval pause as a special-cased control-flow primitive bolted on top of an ephemeral request handler. It treats it as an ordinary suspended state of a durable, server-tracked task — the same durability substrate used for crash recovery.

That matters for three concrete reasons developers should care about:

  1. You don't need a separate orchestration system. The task_id that identifies your approval chain is the same identity used for lease-based crash recovery. There's no second source of truth to keep synchronized.
  2. The wait has no artificial ceiling. Because the chain is durable and not tied to a live process or open connection, "wait for approval" can mean seconds or it can mean a week over a holiday, with identical code.
  3. It composes with the framework layer. If you're already building on LangGraph or Microsoft Agent Framework over the Responses protocol, the approval interrupt is just another checkpoint boundary that resilient_background=True already knows how to persist and rehydrate.

Getting this pattern right is the difference between an agent that enterprises trust with consequential actions and one that gets restricted to read-only, "suggest but don't act" duty forever.

Core Concepts: Task IDs, Entry Modes, and Suspension

To build an approval step correctly you need to be precise about four concepts in the AgentServer SDK (azure-ai-agentserver-core ≥ 2.0.0 for Python, Azure.AI.AgentServer.Core ≥ 1.0.0-beta.28 for .NET, both currently preview surfaces subject to change):

@multi_turn_task is a decorator that turns an async handler into a durable conversation chain. Unlike a one-shot @task (input in, output out, done), a multi-turn task doesn't terminate when the handler returns — it transitions into a suspended state and stays alive under a single task_id until either a new turn arrives or you explicitly delete it.

task_id is the durable work identity that scopes the whole chain. It's caller-chosen, not server-generated, which is the detail that makes human-in-the-loop possible: your application decides the identity up front ("exp-42" for an expense report, a conversation thread ID, a ticket number), and every subsequent turn — including the human's reply, arriving possibly days later from a completely different process — reenters the same chain by reusing that same string.

entry_mode on TaskContext tells your handler why it's being invoked right now. There are three values:

  • fresh — first execution for this (task_id, input_id) pair.
  • resumed — a subsequent turn on an existing chain (this is what fires when the human's decision comes back in).
  • recovered — the container crashed mid-attempt in a previous lifetime and the framework is re-invoking the same attempt from persisted input, without your explicit involvement.

This three-way split is the crux of the whole pattern. Your handler branches on entry_mode to decide whether it's starting fresh, picking up a human decision, or being silently retried after an infrastructure hiccup. Critically, resumed and recovered are different things: resumed is an intentional new turn (the human replied), while recovered is the framework protecting you from a crash that happened before your fresh or resumed turn even finished.

Long-Running Task Entry Modes state machine: fresh leads to suspended, which leads to resumed, which leads to completed, with a recovered branch handling crash re-invocation
Entry modes govern how the framework re-enters your handler: fresh for the first execution, resumed for a genuine next turn (human decision or scheduled check), and recovered when a crash interrupts an attempt before it finishes.

ctx.metadata is small, durable key-value state attached to the task that survives the suspension. The documentation is explicit that this should hold only small references — an expense ID, a step counter — not full payloads. The full request history and generated artifacts belong in your own storage or a FoundryStateStore-backed checkpoint (see Day 1 for the checkpoint/watermark pattern in depth).

The diagram below shows how these pieces fit together end to end.

Architecture diagram of the human-in-the-loop approval flow: an Agent Container's multi_turn_task handler sends a request to TaskManager, which persists SUSPENDED state to FoundryStateStore while a Human Approver reviews and replies, resuming the chain
The suspended chain lives in the durable state store, not in a live process — the container that handled turn 1 can exit entirely before a human ever replies.

Architecture: Where the Approval State Actually Lives

It's worth being explicit about what's happening at the infrastructure level, because "it just suspends" hides a few architectural decisions that matter once you're debugging a stuck task in production.

When a @multi_turn_task handler returns without raising, the TaskManager — a server-side component that Foundry's Agent Service constructs when you call set_resilient_tasks_enabled(True) before host startup — writes the chain's current state to the durable state store and transitions its status to Suspended. This is not an in-memory pause. The container that handled turn 1 can be killed, scaled to zero, or replaced by a new revision entirely, and turn 2 can be served by a completely different container instance, because nothing about the suspension depends on process memory. The only thing that has to survive is the record in the state store and (if you're using framework-level checkpointing) the serialized framework state you wrote there yourself.

This has a direct, useful consequence: an approval wait is not a container-hours cost. A suspended multi_turn_task isn't a thread blocked on input(), and it isn't a container kept warm waiting for a callback. The container that ran turn 1 can exit completely. Whatever compute picks up turn 2 — hours or days later — is a fresh invocation against the same task_id, and the framework's job is purely to route it to the right handler with the right persisted context, not to keep a process alive across the gap.

The TaskStatus enum reflects this lifecycle explicitly: Pending → InProgress → Suspended → Completed. A chain sitting in Suspended is a durable database row (conceptually), not a live process. When you eventually call await approve.delete("exp-42"), you're deleting that durable record — worth noting because the framework does not garbage-collect suspended multi-turn chains automatically the way it cleans up completed one-shot @task records. If your approval chains never get an explicit resolution (the approver never replies, the ticket gets abandoned), you will accumulate orphaned Suspended records unless you build a reaper.

Implementing an Approval Turn Step by Step

Here's a complete, realistic approval handler for an expense-report scenario, annotated beyond what the reference docs show, including the parts most tutorials skip: input validation on resume, and defending against a stale or replayed decision.

from datetime import timedelta
from azure.ai.agentserver.core.tasks import (
    multi_turn_task,
    TaskContext,
    RetryPolicy,
    set_resilient_tasks_enabled,
)

# Must be called once, before host startup, so the TaskManager is
# constructed and the crash-recovery scan runs on container boot.
set_resilient_tasks_enabled(True)


@multi_turn_task(
    name="expense-approval",
    timeout=timedelta(days=7),          # give approvers a realistic window
    retry=RetryPolicy(max_attempts=3),  # applies to handler failures, NOT to the human wait
)
async def approve(ctx: TaskContext[dict]) -> dict:
    if ctx.entry_mode == "resumed":
        # A human decision has arrived on the same task_id.
        decision = ctx.input.get("decision")
        expense_id = ctx.metadata.get("expense_id")

        if expense_id is None:
            # Defensive: metadata should always be set on the fresh turn.
            # If it's missing, the chain state is corrupt — fail loudly
            # rather than silently submitting an unknown expense.
            raise ValueError("Resumed approval turn is missing expense_id metadata")

        if decision not in ("approved", "rejected"):
            # Reject malformed input instead of treating it as ambiguous approval.
            return {"status": "invalid_decision", "expense_id": expense_id}

        if decision == "approved":
            await submit_expense(expense_id)
            return {"status": "submitted", "expense_id": expense_id}
        return {"status": "rejected", "expense_id": expense_id}

    # entry_mode == "fresh": build the request and suspend for a decision.
    expense = await build_expense(ctx.input)
    ctx.metadata["expense_id"] = expense.id
    ctx.metadata["requested_amount"] = expense.amount
    ctx.metadata["requested_at"] = expense.created_at.isoformat()

    await notify_approver(expense)  # see next section

    return {"status": "awaiting_approval", "summary": expense.summary, "expense_id": expense.id}
Enter fullscreen mode Exit fullscreen mode

Two details in this code deserve callouts because they aren't obvious from the quickstart-level docs:

Validate decision on resume. The suspended chain is a durable record that anyone with the right permissions and the right task_id can post a "turn 2" against. Don't assume the resumed input matches the shape you expect — treat it with the same skepticism you'd apply to any external API input, because in a multi-app-surface deployment (Teams bot + web portal + email reply parser all driving the same task_id), it usually is external input.

Set a timeout that matches the real-world wait, not the code's execution time. The timeout on @multi_turn_task bounds the whole chain's lifetime, not a single turn's handler execution. If your approval SLA is "respond within a week," set timedelta(days=7), or the chain will be forcibly failed while still legitimately waiting on a human.

Driving this from your application (a web backend, a bot, a CLI) looks like this:

# Turn 1 — the agent produces a request and the chain suspends.
r1 = await approve.run(task_id="exp-42", input={"amount": 1200, "category": "travel"})
assert r1["status"] == "awaiting_approval"
# Show r1["summary"] to a human through whatever channel you use — Teams
# adaptive card, email, ServiceNow ticket — and wait for their reply.
# This can happen minutes or days later, from an entirely different process.

# Turn 2 — same task_id resumes the suspended chain.
r2 = await approve.run(task_id="exp-42", input={"decision": "approved"})
assert r2["status"] == "submitted"
Enter fullscreen mode Exit fullscreen mode

Note that approve.run() is a plain async call from your application code — there's no special "resume API" distinct from the normal task invocation surface. The framework figures out from the existing Suspended record on "exp-42" that this is a resume, not a fresh start, and sets entry_mode accordingly.

Wiring the Human Side: Notifications and Resume Triggers

The Foundry docs deliberately don't prescribe how the human finds out there's something to approve — that's your application's job, and it's worth being deliberate about because it's the part most likely to silently fail in production. A notify_approver() call that fires-and-forgets to an email API with no delivery confirmation means your "durable" approval chain is only as durable as an SMTP send that nobody checked.

A production-grade notification path typically looks like:

async def notify_approver(expense) -> None:
    # Write an approval record the UI can query, independent of the
    # suspended task itself — this is your operational visibility layer.
    await approvals_table.upsert({
        "task_id": f"exp-{expense.id}",
        "status": "pending",
        "amount": expense.amount,
        "requested_at": expense.created_at.isoformat(),
        "approver_group": expense.approver_group,
    })

    # Fan out to a channel with delivery guarantees, not best-effort webhook.
    await teams_adaptive_card.send(
        channel=expense.approver_group,
        card=build_expense_card(expense),
        # The card's Approve/Reject buttons post back to your own
        # endpoint, which calls approve.run(task_id=..., input={"decision": ...}).
    )
Enter fullscreen mode Exit fullscreen mode

The key architectural point: keep a separate, queryable "pending approvals" projection outside the suspended task record. The task's Suspended state is durable but it's not designed to be a worklist UI's backing store — you generally can't run "give me all suspended expense-approval chains older than 48 hours with no reminder sent" as a query against the task subsystem itself. Maintain that index yourself, keyed by the same task_id, and use it to drive reminders, SLA escalation, and dashboards, while the actual approve/reject decision still flows through approve.run() against the canonical task_id.

Framework Interrupts: LangGraph and Microsoft Agent Framework

If you're not writing raw @multi_turn_task handlers but building on an orchestration framework over the Responses protocol, Foundry's guidance is to use the framework's own interrupt mechanism rather than reimplementing pause/resume by hand — and then make sure the surrounding response stays resilient across the interrupt boundary.

Concretely, for LangGraph this means using interrupt() inside a node and configuring the graph's checkpointer to serialize into Foundry's state store, so that when the Responses host reinvokes your handler after a crash (resilient_background=True), LangGraph's own recovery — not yours — rebuilds the graph from the last checkpoint:

from azure.ai.agentserver.responses import ResponsesAgentServerHost, ResponsesServerOptions

app = ResponsesAgentServerHost(
    options=ResponsesServerOptions(resilient_background=True),
)

@app.response_handler
async def handler(request, context, cancellation_signal):
    if context.is_recovery:
        # Framework-level recovery: rebuild the LangGraph run from its
        # own checkpoint rather than restarting the conversation.
        graph_state = await load_langgraph_checkpoint(context.conversation_id)
        # ... resume graph.stream(..., config={"configurable": {"thread_id": ...}})
    else:
        # Normal path: graph runs and may itself call interrupt() at an
        # approval boundary, which suspends the underlying response.
        ...
Enter fullscreen mode Exit fullscreen mode

For Microsoft Agent Framework, the analogous primitives are RequestInfoEvent and ApprovalRequiredAIFunction, which we introduced in Day 7's piece on the code-first orchestration migration. The relationship between that pattern and what's covered here is important to get straight: ApprovalRequiredAIFunction is a workflow-level declaration that a given tool call requires human sign-off before execution, wired into Agent Framework's SequentialBuilder / HandoffBuilder orchestration graph. The @multi_turn_task approach covered in this article is the lower-level runtime primitive those framework features are ultimately built on when hosted inside Foundry Agent Service. If you're using Agent Framework already, reach for ApprovalRequiredAIFunction first — it gives you the same suspend/resume durability with considerably less hand-written plumbing. Reach for raw @multi_turn_task when you're not using an orchestration framework at all, or when you need approval semantics the framework's built-in primitive doesn't expose (multi-approver quorum, conditional auto-approval thresholds, cross-task approval batching).

A Real-World Scenario: Expense Approval With Escalation

Let's extend the expense example to something closer to what a real finance-ops team would actually deploy, adding a second wrinkle: escalation if nobody responds within 48 hours.

from datetime import timedelta, datetime, timezone

@multi_turn_task(name="expense-approval-v2", timeout=timedelta(days=7))
async def approve_v2(ctx: TaskContext[dict]) -> dict:
    if ctx.entry_mode == "resumed":
        kind = ctx.input.get("kind")

        if kind == "escalation_check":
            # A scheduled external trigger (not a human) posts this turn
            # periodically to ask "has anyone approved yet?"
            if datetime.now(timezone.utc) > _deadline(ctx.metadata):
                await escalate_to_manager(ctx.metadata["expense_id"])
                ctx.metadata["escalated"] = True
            # Re-suspend: escalation checks don't resolve the chain.
            return {"status": "awaiting_approval", "escalated": ctx.metadata.get("escalated", False)}

        if kind == "decision":
            decision = ctx.input.get("decision")
            expense_id = ctx.metadata["expense_id"]
            if decision == "approved":
                await submit_expense(expense_id)
                return {"status": "submitted", "expense_id": expense_id}
            return {"status": "rejected", "expense_id": expense_id}

        return {"status": "ignored_unknown_input"}

    # Fresh turn.
    expense = await build_expense(ctx.input)
    ctx.metadata["expense_id"] = expense.id
    ctx.metadata["requested_at"] = datetime.now(timezone.utc).isoformat()
    await notify_approver(expense)
    await schedule_escalation_check(ctx.task_id, delay=timedelta(hours=48))
    return {"status": "awaiting_approval", "expense_id": expense.id}
Enter fullscreen mode Exit fullscreen mode

This pattern — using a scheduled resume turn (kind: "escalation_check") rather than only human-driven ones — is a useful generalization worth internalizing: nothing says a "resumed" turn has to come from a person. Any external trigger (a cron job, a Logic App timer, another agent) that posts to the same task_id is a legitimate way to drive the chain forward, as long as your handler discriminates between input shapes (kind field above) so an escalation check doesn't get mistaken for an approval decision. This is exactly the kind of input-validation discipline flagged earlier — it stops being optional the moment more than one caller can legitimately post to a suspended chain.

Production Considerations

Set realistic per-chain timeouts, and monitor what falls outside them. A timeout on a multi_turn_task fails the whole chain if it's exceeded, including the suspended wait. Whatever your organizational SLA is for approvals (24 hours, a week, "until the next board meeting"), set the timeout to comfortably exceed it, and separately alert on approvals that are taking unusually long — that's an operational signal (an approver is unreachable, a request is ambiguous), not a framework failure.

Build the reaper. Suspended multi-turn chains are not automatically cleaned up. If your business process allows an approval request to simply be abandoned (the requester cancels the expense, the ticket gets closed elsewhere), you need your own job that finds stale Suspended records via your parallel "pending approvals" index and calls .delete() explicitly, or you'll accumulate state-store bloat indefinitely.

Idempotency on the resume path matters more than on the fresh path. Approval UIs (Teams cards, email links) are notoriously prone to double-submission — a user double-clicks "Approve," or a webhook retries after a timeout even though the first call succeeded. Use if_last_input_id on the resume call where your driving application can track it, and make submit_expense() itself idempotent (keyed by expense_id, not by task invocation) so a duplicate "approved" turn doesn't double-submit the underlying business action.

Decide where the audit trail lives. Regulatory and compliance requirements around "who approved what and when" are common in exactly the domains (expense, procurement, infra changes) where this pattern is most useful. The task's own state transitions are not automatically a compliant audit log — capture the decision, the identity of the approver, and the timestamp explicitly into your own storage (or a compliance-grade log sink) at the moment you construct the resume input, not after the fact.

Security Considerations

Anyone who can call .run() with the right task_id can post a "decision." The task subsystem's job is durability and state management, not authorization. It is entirely your application's responsibility to verify that the caller posting {"decision": "approved"} against "exp-42" is actually a member of the approver group for that expense, has an active session, and isn't replaying a stale link. Do this check in the code that calls approve.run(), before the resume even happens — not inside the handler, where by the time you've read ctx.input the framework has already accepted the turn as a legitimate resume of the chain.

Guard against task_id enumeration and prediction. If your task_id scheme is guessable ("exp-42", "exp-43"...), a malicious actor could attempt to post decisions against IDs they don't own. Prefer a task_id that embeds a non-guessable component (a UUID segment, an HMAC of the internal record ID) when the chain governs a consequential action, and always cross-check the caller's identity against the resource the task_id represents, not just against the string itself.

Treat ctx.metadata as durable but not secret. It's server-backed state, not a secrets vault. Don't stash API keys, tokens, or PII beyond what's operationally necessary in metadata — remember the guidance to keep only small references there; that's a security boundary as much as a size-limit one.

Approval requests are a prompt-injection surface too. If the "expense summary" or any agent-generated content shown to the human approver is itself derived from untrusted input (a user-submitted expense description, a scraped web page), sanitize what's rendered in the approval card. An approver clicking "Approve" on a card whose displayed text was manipulated by injected content is a real failure mode, not a theoretical one — treat the approval UI with the same rendering hygiene you'd apply to any user-generated content surface.

Performance and Scale

The good news architecturally: suspended chains cost you nothing in compute while they wait. There's no polling loop, no kept-alive connection, no container reserved per pending approval. This means the pattern scales to tens of thousands of concurrently pending approvals without any capacity planning beyond the state store's own storage and throughput limits — a state store item caps at 1 MB of serialized JSON, and a store name (which you choose to scope by workflow/session/thread) can be 1–128 characters, with up to 16 tags per item, all of which are generous for approval metadata.

The place scale does bite is your notification and resume fan-in path, not the task subsystem itself. If you're driving thousands of resumes per minute through a webhook handler that calls approve.run() synchronously, that handler is an ordinary web endpoint subject to ordinary web-endpoint scaling concerns — put it behind whatever autoscaling and queuing discipline you'd apply to any high-throughput API, independent of the Foundry task layer underneath it.

Watch TaskConflictError: a non-steerable task rejects a concurrent start against an in-flight task_id. If two processes race to resume the same approval (a double-click plus a webhook retry landing at the same moment), one will get this exception — treat it as an expected, retryable-with-backoff condition in your resume-driving code, not an unhandled error.

Cost Considerations

The direct cost surface of the approval pattern itself is small: state-store storage for the suspended record and its metadata, for however long the chain remains pending, plus whatever notification channel you use (Teams, email, SMS) at its own pricing. There's no token cost for the wait itself, since no model call happens between suspension and resume.

The cost trap to watch is indirect: if your fresh turn does expensive work before the approval boundary — a large document analysis, several tool calls, an LLM call to generate the approval summary — and your escalation or reminder logic causes the whole chain to be retried from fresh rather than resumed cleanly, you'll pay for that expensive work repeatedly. This is why the entry-mode discrimination matters as much for cost control as for correctness: a handler that doesn't check entry_mode and unconditionally rebuilds the expense summary on every turn — including escalation checks — will silently re-run an LLM call on every 48-hour escalation ping, indefinitely, for a chain that might sit pending for a week.

Common Mistakes and Pitfalls

Forgetting set_resilient_tasks_enabled(True). Declaring @multi_turn_task does not, by itself, activate the resilient task subsystem. Without the explicit opt-in call before host startup, get_task_manager() raises TaskManagerNotInitialized, and neither .run() nor .start() works at all. This is the single most common "it just doesn't suspend" bug reported against early previews of this feature.

Conflating resumed with "the human approved." A resumed entry mode only means a new turn arrived on an existing chain — it says nothing about what that turn contains. Escalation pings, reminder acknowledgments, and actual approve/reject decisions are all resumed turns. Always discriminate on the input payload's own shape, not on entry_mode alone.

Storing the full request payload in ctx.metadata "for convenience." The documented guidance is explicit: metadata is for small references, not full history. Beyond the practical 1 MB item ceiling, doing this couples your durable identity state to your business payload schema in a way that makes future migrations painful. Use FoundryStateStore checkpoints (see Day 1) for anything beyond a handful of scalar fields.

Treating the suspended wait as bounded by the handler's own retry policy. RetryPolicy governs retries of handler failures; it has nothing to do with how long a chain can sit Suspended waiting for a human. That's governed by timeout, a separate and easily overlooked parameter.

Not building the reaper. Covered above under production, but worth restating as a pitfall: teams that ship this pattern to production without a cleanup job for abandoned approval chains eventually notice unbounded growth in their state store and have no clean way to distinguish "still legitimately pending" from "abandoned two months ago" without one.

Alternatives and Trade-offs

Rolling your own with Durable Functions or Temporal. You get more control over workflow visualization and cross-cutting concerns (Temporal's UI, for instance, is genuinely good at surfacing stuck workflows) at the cost of running and paying for a separate orchestration system alongside Foundry Agent Service, and manually bridging conversational/agent state between the two. If you're deeply invested in one of these already for non-agent workflows, integrating rather than replacing may be the pragmatic choice — but for agent-native approval flows, the native multi_turn_task primitive removes an entire system from your architecture.

Synchronous polling with a client-side wait loop. Simpler to reason about for short waits (seconds to a couple of minutes) where you can afford to hold a connection or poll a status endpoint, but it does not scale to human-timescale waits (hours to days) without building your own durability layer — which is exactly what you'd be reimplementing.

Framework-native interrupts (LangGraph interrupt(), Agent Framework ApprovalRequiredAIFunction). As discussed above, prefer these when you're already using the framework, since they compose with the resilient Responses layer with less code. The trade-off is less flexibility for approval patterns the framework didn't anticipate (multi-approver quorum, conditional thresholds), where dropping to raw @multi_turn_task gives you full control.

Practical Recommendations

  • Default to framework-native interrupts (Agent Framework ApprovalRequiredAIFunction, LangGraph interrupt()) if you're already orchestrating with one of those frameworks; reach for raw @multi_turn_task when you need custom approval semantics or aren't using an orchestration framework.
  • Always validate the shape and authorization of resumed input — never trust that a resumed turn is the decision you expect.
  • Maintain a separate, queryable "pending approvals" projection for operational visibility; don't rely on the task subsystem as your worklist backing store.
  • Set timeout to match your real-world SLA, not your code's execution time, and alert separately on approvals exceeding expected wait times.
  • Build the reaper job before you ship, not after you notice state-store growth in production.
  • Keep ctx.metadata to small references; put anything substantial into your own storage or a FoundryStateStore checkpoint.

Conclusion

Human-in-the-loop approval in Microsoft Foundry isn't a bolted-on feature — it's what a multi_turn_task chain already does when a turn suspends and waits for the next input, whether that input is a human's decision, a scheduled escalation check, or a framework-level interrupt resuming from a checkpoint. The durability guarantees that make long-running agents crash-resilient (Day 1's leases, checkpoints, and entry_mode recovery) are the same guarantees that let an approval wait stretch from ninety seconds to nine days without your application needing a second orchestration system to track it. The engineering work that remains is squarely on your side of the boundary: validating resumed input rigorously, keeping an operational index of pending approvals, authorizing who's allowed to resume a chain, and building the cleanup job for the approvals nobody ever answers. Get those right, and you have a pattern that lets an enterprise finally trust an agent to propose a consequential action and wait, correctly, for a person to say yes.

If you're building an agent that touches money, infrastructure, or anything requiring sign-off, start with the @multi_turn_task primitive covered here before you reach for a separate workflow engine — you likely already have what you need.

References

(Note: APIs referenced are in preview and subject to change; verify current package versions and behavior against Microsoft Learn before shipping to production.)

Top comments (0)