DEV Community

Cover image for Routing OpenAI Computer-Use & Browser Use Agent Escalations to Human Sales Teams in Real Time
InstaChime
InstaChime

Posted on Originally published at instachime.com Fully Autonomous

Routing OpenAI Computer-Use & Browser Use Agent Escalations to Human Sales Teams in Real Time

Browser agents have moved from demos to production tooling. Agents can now read a web page, fill in multi-page forms, configure software and carry out purchase-intent tasks through the same graphical interface a person would use. They still reach points where a human has to take over: a custom-quote form, a high-value order, a login that needs two-factor authentication, or a page layout the agent can’t parse.

For sales and revenue operations teams, the practical question is what happens at that moment. Speed matters. A widely cited Harvard Business Review analysis of about 1.25 million inbound leads (Oldroyd, McElheran and Elkington, 2011) found that firms contacting a lead within an hour were nearly seven times as likely to qualify it as firms that waited one more hour. They were more than 60 times as likely as firms that waited 24 hours or longer. That study covered web-form leads, not agent-generated ones, but the lesson carries over: if an agent surfaces a high-intent moment, a human should see it while it is still live.

This guide shows how to build a real-time escalation pipeline. When an agent you operate hits a trigger, the pipeline posts an interactive “claim” card to Slack or Microsoft Teams. One rep claims the lead atomically, and that rep takes over the live browser session.

A note on OpenAI’s product names. OpenAI’s Operator, the product that popularized this pattern, no longer exists. OpenAI folded it into ChatGPT agent in July 2025, and the standalone site was shut down on August 31, 2025. OpenAI’s help center now says ChatGPT agent is itself no longer available. It points users to ChatGPT Work and to a cloud-browser feature in ChatGPT instead. The computer-use capability remains available to developers through the OpenAI API, so that is what this guide builds on.


1. The Agentic Web Navigation Landscape

OpenAI: from Operator to the Computer Use API

Operator launched on January 23, 2025, powered by a model called Computer-Using Agent (CUA). CUA combined GPT-4o’s vision with reinforcement-learning-based reasoning. It worked from screenshots and a virtual mouse and keyboard. It did not parse the page’s DOM and needed no custom API integrations. At launch OpenAI reported 38.1% on OSWorld, 58.1% on WebArena and 87% on WebVoyager. It also noted that CUA still had to improve on harder benchmarks like WebArena before matching human performance.

Today developers reach this capability through the Responses API. OpenAI’s current documentation describes two integration styles:

  • Code execution. The model writes code (for example with Playwright or PyAutoGUI), and your application runs it in an isolated browser or desktop. OpenAI recommends this for its newest models.
  • The computer tool. The model returns structured mouse and keyboard actions, which your application executes before sending back a screenshot.

In both styles you provide and operate the environment. OpenAI’s docs also say that continuing an API response does not restore a browser session, login state or runtime variables, so the browser must be kept alive separately. This matters a great deal for human takeover (see Section 6).

Open source: browser-use

browser-use is an MIT-licensed Python library (Python 3.11 or later) that lets an LLM drive a browser to complete tasks such as filling forms and extracting data. It supports several model providers, including OpenAI, Anthropic and Google models, plus its own hosted ChatBrowserUse models. The project originally used Playwright. In August 2025 it moved to talking to Chrome directly over the Chrome DevTools Protocol (CDP). A hosted cloud agent is also available for production use.

Custom actions that the agent can call are registered with @tools.action(...) on a Tools object. The older name for this was @controller.action(...) on a Controller.

Why autonomous agents still need human escalation

Benchmark results are not a guarantee of reliability. OpenAI’s own guidance for computer use is to confirm consequential actions with a person and to verify the actual outcome rather than trusting the model’s final message. In a sales context, typical escalation triggers include:

  • High-value thresholds. The agent reaches a quote or configuration above a deal size you define, for example an annual contract value above your enterprise-review threshold.
  • Custom-quote or negotiation pages. There are no fixed options for the agent to select.
  • Authentication and verification walls. These include two-factor prompts, CAPTCHAs and credential hand-offs. Browserbase’s documentation lists two-step verification as a case where a human should take over through a live view.
  • UI drift. A site change breaks the agent’s expected path and it stalls or loops.

This guide assumes you operate the agent, such as an SDR or procurement agent, a quoting assistant, or a buyer-side agent you built. If a prospect’s own third-party agent visits your site, you cannot hook its halt events. In that case the trigger has to come from signals on your side, such as pricing-page or quote-form events.


2. Designing the Real-Time, Event-Driven Escalation Architecture

When a trigger fires, the system moves from autonomous mode to assisted mode. The agent stops acting, the browser session stays alive, and a notification goes out.

+-------------------------------------------------------------+
|                 Autonomous Browser Agent                    |
|   (OpenAI computer-use harness or `browser-use` runner)     |
+-------------------------------------------------------------+
                               |
                   [Escalation trigger fires]
        (high-value quote, custom-quote form, stall, 2FA)
                               |
                               v
+-------------------------------------------------------------+
|              Orchestration Backend (FastAPI / Node.js)      |
|  - Stores escalation record (reason, URL, form state)       |
|  - Captures screenshot, redacts PII                         |
|  - Holds the live-view URL server-side (not in the card)    |
+-------------------------------------------------------------+
             /                                   \
            /          (Parallel dispatch)         \
           v                                       v
+-----------------------+               +-----------------------+
|     Slack app         |               |    Teams bot          |
|  - Block Kit message  |               |  - Adaptive Card      |
+-----------------------+               +-----------------------+
           \                                       /
            \                                     /
             v                                   v
+-------------------------------------------------------------+
|                  Human Sales Representative                 |
|      Clicks "Claim lead" -> atomic claim -> live view link  |
+-------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Core components

  1. Agent runner and state store. A backend service (typically Python with FastAPI, or Node.js) runs the agent loop. Escalation state, such as pending, claimed or released, lives in a store such as Redis.
  2. Escalation endpoint. Your agent’s custom tool or harness posts a structured event to this endpoint. It is your own endpoint, not an OpenAI webhook.
  3. Notification dispatcher. It turns the event into a Slack Block Kit message and a Teams Adaptive Card, each with a screenshot and the key fields.
  4. Interaction handlers. Slack requires your app to acknowledge every interaction payload within three seconds. Do the slow work (database writes, CRM calls) after acknowledging.
  5. Takeover layer. A cloud browser provider with a human-usable live view, such as Browserbase or Steel, or a self-hosted browser exposed over CDP. It lets the rep see and control the same browser the agent was using.

3. Implementing the Escalation Trigger with browser-use

The example below runs the agent on a remote Browserbase session and defines a custom tool that escalates when the agent finds a quote form. The domain, company and numbers are illustrative.

import asyncio
import os

import httpx
from browser_use import ActionResult, Agent, Browser, ChatOpenAI, Tools
from browserbase import Browserbase

# A remote browser session that a human can later take over.
bb = Browserbase(api_key=os.environ["BROWSERBASE_API_KEY"])
session = bb.sessions.create(
    project_id=os.environ["BROWSERBASE_PROJECT_ID"],
    timeout=1800,  # seconds; leave headroom for a human takeover
)
live_view_url = bb.sessions.debug(session.id).debuggerFullscreenUrl

tools = Tools()


@tools.action(
    description=(
        "Escalate to the human sales team. Use when the page shows a custom-quote "
        "form, enterprise pricing without fixed options, or a quote above the "
        "enterprise review threshold. After calling this, stop acting."
    )
)
async def escalate_to_sales(
    company_name: str, estimated_seats: int, page_url: str, quoted_price: str
) -> ActionResult:
    payload = {
        "event_type": "high_intent_escalation",
        "session_id": session.id,
        "company_name": company_name,
        "estimated_seats": estimated_seats,
        "page_url": page_url,
        "quoted_price": quoted_price,
        "live_view_url": live_view_url,  # stays server-side; never put in the channel card
        "status": "pending_claim",
    }
    try:
        async with httpx.AsyncClient(timeout=5) as client:
            resp = await client.post(os.environ["ESCALATION_WEBHOOK_URL"], json=payload)
            resp.raise_for_status()
    except httpx.HTTPError as exc:
        return ActionResult(
            extracted_content=f"Escalation failed: {exc}. Retry once, then stop and report."
        )
    return ActionResult(
        is_done=True,
        success=True,
        extracted_content="Escalated to the human sales team. Stopping here.",
    )


async def main():
    agent = Agent(
        task=(
            "Go to https://example.com/pricing, evaluate the enterprise tier for a "
            "250-seat team, and if a custom-quote form appears, call the escalation tool."
        ),
        llm=ChatOpenAI(model="gpt-5.5"),
        browser=Browser(cdp_url=session.connect_url),
        tools=tools,
    )
    await agent.run()


asyncio.run(main())
Enter fullscreen mode Exit fullscreen mode

A few notes on this pattern:

  • Ending the run is not ending the session. The agent’s run completes when the escalation tool returns is_done=True, but the remote browser must stay open for the rep. Set a session timeout long enough for a takeover, and check your browser-use version’s keep-alive behavior so the client doesn’t close the browser when the run ends.
  • The model fills in the tool’s arguments. Treat everything the agent extracts from a web page as untrusted. OpenAI’s computer-use guidance says page text cannot grant permission or override the user’s instructions. Validate and escape these fields before they reach a team channel (see Section 7).
  • If you use OpenAI’s API directly, the same idea applies. Expose an escalation function tool alongside the computer tool or your code-execution tool. When the model calls it, stop the loop and post the event instead of continuing.

4. Crafting Interactive Claiming Cards for Slack and Microsoft Teams

When the backend receives an escalation event, it posts a card to the sales channel. The common rule for high-intent leads is that the first rep to claim owns the lead. Round-robin assignment with a fallback timer is a common alternative. Either way, an interactive card prevents duplicate outreach and gives the rep context at a glance.

Slack Block Kit

Slack’s Block Kit supports headers, field sections, images and button actions. Two practical notes:

  • The image block needs a URL that Slack’s servers can fetch. Use a short-lived signed URL for a redacted screenshot rather than a permanent public link.
  • The card deliberately has no “spectate” link. A live-view URL gives control of a real browser session, and Steel’s documentation, for one, states that anyone with a debug URL can view or interact with the session. Send the link only to the rep who claims the lead.
{
  "channel": "C0123456789",
  "text": "High-intent agent escalation: Acme Global Corp",
  "blocks": [
    {
      "type": "header",
      "text": { "type": "plain_text", "text": "🚨 High-intent agent escalation", "emoji": true }
    },
    {
      "type": "section",
      "fields": [
        { "type": "mrkdwn", "text": "*Company:*\nAcme Global Corp" },
        { "type": "mrkdwn", "text": "*Estimated value:*\n$65,000 ACV (250 seats)" },
        { "type": "mrkdwn", "text": "*Agent state:*\nStopped at custom-quote form" },
        { "type": "mrkdwn", "text": "*Page:*\n<https://example.com/enterprise|View target page>" }
      ]
    },
    {
      "type": "image",
      "image_url": "https://assets.example.com/escalations/esc_8f3a/screenshot.png?sig=EXAMPLE",
      "alt_text": "Browser screenshot at the moment of escalation"
    },
    {
      "type": "actions",
      "elements": [
        {
          "type": "button",
          "text": { "type": "plain_text", "text": "🎯 Claim lead & take over", "emoji": true },
          "style": "primary",
          "value": "esc_8f3a",
          "action_id": "sales_claim_lead"
        }
      ]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

With Slack’s Bolt framework for Python, the click handler acknowledges first, then claims and updates the message:

@app.action("sales_claim_lead")
def handle_claim(ack, body, client):
    ack()  # Slack requires an acknowledgment within 3 seconds
    escalation_id = body["actions"][0]["value"]
    rep_id = body["user"]["id"]
    channel, ts = body["channel"]["id"], body["message"]["ts"]

    won, owner = try_claim(escalation_id, rep_id)  # see Section 5
    if won:
        client.chat_update(channel=channel, ts=ts, text=f"Claimed by <@{rep_id}>",
                           blocks=claimed_blocks(escalation_id, rep_id))
        send_takeover_link_privately(rep_id, escalation_id)  # DM or ephemeral message
    else:
        client.chat_postEphemeral(channel=channel, user=rep_id,
                                  text=f"Already claimed by <@{owner}>.")
Enter fullscreen mode Exit fullscreen mode

The interaction payload also carries a response_url. You can POST to it to update the original message, up to five times within 30 minutes of receiving the payload. Non-ephemeral messages can alternatively be updated with chat.update, as above.

Microsoft Teams Adaptive Cards

Teams renders Adaptive Cards. For a claim button that works across many reps, use the Universal Actions model: Action.Execute sends the click to your bot, and your bot can respond with an updated card. Some constraints:

  • Card version. Microsoft’s Teams documentation says Universal Actions need Adaptive Card schema version 1.5 or later. (Some Adaptive Cards docs say 1.4, so use 1.5 to be safe.)
  • Bot required. Action.Execute is handled by a bot. Your bot must handle the resulting adaptiveCard/action invoke request.
  • Fallback. Microsoft recommends adding Action.Submit as a fallback so the card still works on clients that don’t render Action.Execute.
{
  "type": "AdaptiveCard",
  "$schema": "http://adaptivecards.io/schemas/adaptive-card.json",
  "version": "1.5",
  "body": [
    {
      "type": "TextBlock",
      "text": "🚨 Agent escalation: human sales rep needed",
      "weight": "Bolder",
      "size": "Medium",
      "color": "Attention"
    },
    {
      "type": "FactSet",
      "facts": [
        { "title": "Prospect:", "value": "Acme Global Corp" },
        { "title": "Intent signal:", "value": "250-seat enterprise quote form" },
        { "title": "Trigger:", "value": "Custom pricing required" }
      ]
    },
    {
      "type": "Image",
      "url": "https://assets.example.com/escalations/esc_8f3a/screenshot.png?sig=EXAMPLE",
      "size": "Large",
      "altText": "Browser screenshot at the moment of escalation"
    }
  ],
  "actions": [
    {
      "type": "Action.Execute",
      "title": "Claim lead & take over",
      "verb": "claim_lead",
      "data": { "escalationId": "esc_8f3a" },
      "fallback": {
        "type": "Action.Submit",
        "title": "Claim lead & take over",
        "data": { "escalationId": "esc_8f3a", "verb": "claim_lead" }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

As in Slack, send the live-view link to the rep who claims the lead, for example in a one-to-one bot message. Don’t put it on the shared card.


5. Managing State, Concurrency and the “First to Claim” Race Condition

With a dozen reps watching one channel, two people can click “Claim” within milliseconds of each other. The claim must be a single atomic operation. In Redis, SET key value NX EX seconds does this: it sets the key only if it doesn’t exist and applies a TTL in the same step. Avoid a separate “check, then set” sequence. The older SETNX command can’t set an expiry atomically and is regarded as deprecated in favor of SET ... NX.

import redis

r = redis.Redis(host="localhost", port=6379, db=0, decode_responses=True)
CLAIM_TTL_SECONDS = 15 * 60  # starting point; tune to your handoff process


def try_claim(escalation_id: str, rep_id: str) -> tuple[bool, str]:
    """Return (won, current_owner). Safe to call repeatedly for the same rep."""
    key = f"claim:{escalation_id}"
    for _ in range(2):  # retry once in case the key expires between SET and GET
        if r.set(key, rep_id, nx=True, ex=CLAIM_TTL_SECONDS):
            r.set(f"status:{escalation_id}", "human_active")
            return True, rep_id
        owner = r.get(key)
        if owner == rep_id:
            return True, rep_id  # idempotent: same rep clicked twice
        if owner is not None:
            return False, owner
    return False, "unknown"
Enter fullscreen mode Exit fullscreen mode

Some practical points:

  • Keep handlers idempotent. Platforms retry deliveries that aren’t acknowledged in time, and people double-click.
  • A single Redis key is fine for a single claim. If you need stronger durability guarantees, such as surviving a Redis failover, use a database with a unique constraint or a conditional write instead.
  • Update the card after the claim. Replace the button with “Claimed by …” so no one else clicks it. In Slack use chat.update or response_url. In Teams, return the updated card from your bot’s invoke response.

6. Seamless Handover: From Autonomous Agent to Human Control

Once a rep wins the claim, the system must hand over the exact browser state the agent was using.

  1. Stop the agent. The agent’s loop ends, as in Section 3, and nothing else may click in that session. The state a rep needs, including the page, form fields, cookies and storage, lives in the browser session itself, not in the model conversation. That is why the session must outlive the agent’s run.
  2. Give the rep a live view. Cloud browser providers offer this. Browserbase’s live view is described as an interactive window to display or control a session, with human-in-the-loop uses such as taking control, handling iframes and delegating credentials. You get a link by calling the sessions debug endpoint (debuggerFullscreenUrl). Steel offers embeddable live sessions with an interactive option for remote mouse and keyboard input. With a self-hosted browser, you can expose a remote-debugging connection over CDP behind your own authentication.
  3. Sync to the CRM. Create or update the lead and deal in HubSpot, Salesforce or your CRM of choice, carrying the company, the agent’s extracted fields, the page URL and the escalation reason. A custom “lead source” value such as “AI agent escalation” makes these deals easy to report on later.
  4. Let the rep act. The rep sees what the agent saw. They can finish the form, take a call or chat with the prospect, or complete a verification step, such as a one-time code, themselves.

7. Best Practices for Production

  • Calibrate escalation thresholds. Escalate on clear signals only: a deal-size threshold, an explicit custom-quote path, or a confidence signal from your own logic. Escalating every ambiguity burns out reps and teaches them to ignore the channel.
  • Run two timers, not one. Choose values that fit your team, since these are starting points. If no rep claims an escalation within a couple of minutes, re-notify an on-call rep or manager channel. If a rep claims it but makes no contact within a few minutes, release the claim and re-queue it. Given the lead-response research above, shorter is better. Keep the claim TTL (Section 5) consistent with these timers so they don’t contradict each other.
  • Escape and validate agent-supplied text. Slack requires &, < and > to be escaped in message text. Without that, scraped page content could inject links or mentions such as @channel into your card. Cap field lengths and validate types, since the model produces these values from untrusted pages.
  • Protect the live-view link and screenshot. Treat the live-view URL as a credential. Deliver it only to the claiming rep, give the session a bounded lifetime and revoke it afterwards. Redact payment data, personal data and credentials from screenshots before posting them to a shared channel. Apply data minimization and retention limits to what your pipeline stores. Depending on your users, regulations such as GDPR or CCPA may apply, so involve your privacy or legal team.
  • Bound the agent itself. OpenAI’s guidance is to run computer use in an isolated browser or VM with an allow list of sites and actions, to set step, time or cost limits, and to verify outcomes. Apply the same limits to your escalation tool so a looping agent can’t flood a channel.
  • Log human overrides. Record when reps decline, re-route or correct an escalation. Use that data to refine your trigger rules and prompts and reduce unnecessary escalations.

Conclusion

Browser agents are good at the repetitive parts of web interaction and weak at judgment calls, negotiation and trust-sensitive steps. A well-built escalation pipeline respects that split. Your agent runs on a session a human can later control, a trigger tool posts a structured event, and a Slack or Teams card lets exactly one rep claim the lead atomically. The rep then takes over the live session with full context. The product names around this pattern have shifted quickly, with Operator becoming ChatGPT agent and then giving way to Work and the cloud browser. The architecture stays stable because it depends on a few durable pieces: a browser session you control, a claim lock, and a fast notification path to the right person.


Sources


Originally published at https://instachime.com/blog/routing-openai-computer-use-browser-use-agent-escalations-to-human-sales-teams-in-real-time

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.