Most voice AI quality checks happen after the call ends. A transcript is stored, someone reviews a small sample later, and any mistake has already reached the customer.
What if a supervisor could watch an AI support call while it was happening, send the assistant a private coaching instruction, and join the same conversation when human help was needed?
That is what the Live Support Coach Room example demonstrates. It is a TypeScript application running on Telnyx Edge Compute with:
- a live supervisor dashboard
- one durable room for each AI conversation
- policy-based coaching nudges
- a WebRTC softphone for human escalation
- a simulator that exercises the workflow without placing a real call
This tutorial focuses on the architecture and the fastest way to explore it locally and on Telnyx Edge.
What the finished flow looks like
The application uses two browser tabs:
- The caller simulator creates a conversation and submits caller or assistant turns.
- The supervisor dashboard lists active rooms, streams the transcript, shows policy flags, and exposes an escalation button.
Try entering the same account number twice in the simulator. The room detects that the caller has repeated sensitive account information without identity verification and injects a coaching instruction telling the assistant to verify the caller's date of birth.
The same room can then dial a supervisor and join that person to the active AI conversation.
The high-level flow is:
AI Assistant event stream
|
v
AssistRelay
|
v
CoachRoom for this conversation
| | |
| | +--> Call Control + WebRTC escalation
| +--------------> coaching policy + injected nudge
+-------------------------> supervisor dashboard WebSocket
Why the event stream is a side channel
The assistant's websocket_settings points to the application's /agents/assist endpoint. Telnyx opens one WebSocket for each conversation and sends events such as:
session.created
conversation.item.created
response.text.delta
telnyx.call.answered
telnyx.call.hangup
session.ended
This socket observes the conversation and can inject text turns, but it does not carry the call media and does not replace the model.
That distinction matters. If the coaching application becomes unavailable, the call continues. The caller does not lose audio because the monitoring surface crashed.
The conversation event stream documentation describes it as a live side channel. It is currently beta, and clients should ignore event types they do not recognize because new types may be added.
The stream also is not a guaranteed transcript archive. Events that occur while the socket is unavailable are not replayed. Use it for live supervision, and use the Conversations API after the call when you need a complete system-of-record transcript.
Configure the assistant with the deployed WebSocket URL:
curl -X POST "https://api.telnyx.com/v2/ai/assistants/$ASSISTANT_ID" \
-H "Authorization: Bearer $TELNYX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"websocket_settings": {
"enabled": true,
"url": "wss://live-support-coach-room-<id>.telnyxcompute.com/agents/assist",
"auth_ref": "<your-integration-secret-reference>"
}
}'
auth_ref points to an integration secret. Telnyx resolves it and sends the value as a bearer token when opening the WebSocket.
Route every conversation to its own durable room
The application mounts two Agent SDK WebSocket surfaces under /agents:
const handleAgents = mountAgents<Env>((env) => ({
assist: env.RELAY,
"coach-room": env.COACHROOMS,
}));
/agents/assist receives the assistant event stream. On session.created, the relay uses the conversation ID to locate that conversation's CoachRoom actor.
Each room owns the state for exactly one conversation:
- transcript turns
- policy flags
- nudge count
- stream status
- escalation state
- a durable silence timer
- the final coach log
This is a good fit for a stateful actor: the actor identity is the conversation identity, and related events are handled by the same durable instance.
A separate CoachRegistry actor tracks active and ended rooms for the dashboard picker.
Stream snapshots and live changes to the dashboard
The supervisor connects to:
wss://<function-host>/agents/coach-room/{conversation_id}
When a supervisor opens a room, the dashboard receives its current state and then follows live updates. This means a supervisor joining halfway through a call can see what happened before opening the dashboard instead of starting with an empty screen.
The application also exposes regular HTTP routes:
| Route | Purpose |
|---|---|
GET /rooms |
List active and ended rooms |
GET /rooms/{id}/snapshot |
Read the current room state |
GET /rooms/{id}/log |
Read the room's coach log |
POST /rooms/{id}/join |
Escalate to a human supervisor |
GET /health/liveness |
Basic liveness check |
GET /health/readiness |
Configuration readiness check |
Inject a coaching nudge
The room watches conversation events and evaluates a deliberately small policy. When a rule fires, it sends a conversation.item.create frame back over the assistant stream.
Conceptually, the frame looks like this:
{
"type": "conversation.item.create",
"item": {
"type": "message",
"role": "assistant",
"content": [
{
"type": "input_text",
"text": "Verify the caller's date of birth before sharing account details."
}
]
}
}
The instruction becomes part of the live conversation context, allowing the assistant to change course without ending or restarting the call.
Two settings constrain the policy:
NUDGE_MAX_PER_CALL=3
SILENCE_SECS=90
The first prevents an overactive policy from flooding one conversation. The second schedules a check-in after extended silence. The silence watcher is cancelled when session.ended arrives.
For a production deployment, the simple rule can be replaced with policies for compliance language, authentication, escalation risk, or objection handling.
Bring a human into the AI conversation
When the supervisor clicks Escalate to supervisor, the room:
- creates an outbound call to the configured supervisor device
- joins that call leg to the active AI conversation with
ai_assistant_join - answers the call in the browser through the Telnyx WebRTC SDK
The softphone answers muted and unmutes after the join succeeds. This avoids exposing setup noise to the caller.
The human joins the existing conversation as another participant rather than receiving a cold transfer. Telnyx's multi-participant call documentation covers the broader pattern for adding people to active AI calls.
Run the simulator first
The included simulator is the quickest way to understand the state flow because it does not require telephony.
Clone and install the project:
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/live-support-coach-room
npm install
Create your environment file:
cp .env.example .env
Authenticate and prepare the Edge function:
telnyx-edge auth api-key set <your_telnyx_api_key>
telnyx-edge new-func -l ts -n live-support-coach-room --from-dir .
telnyx-edge types
Run the smoke test:
npx tsx smoke_test.ts
Deploy it:
telnyx-edge ship
Then open these pages from the deployed function:
https://<function-host>/caller
https://<function-host>/dashboard
Start a simulated call in the caller tab. Enter an account number twice, then watch the policy flag and coaching instruction appear in the dashboard.
Configure the live escalation path
The simulated coaching flow works without a call, but live escalation needs the following Edge secrets:
telnyx-edge secrets add TELNYX_API_KEY "<your-key>"
telnyx-edge secrets add COACH_AUTH "<integration-secret-value>"
telnyx-edge secrets add CALL_CONTROL_CONNECTION_ID "<connection-id>"
telnyx-edge secrets add TELNYX_NUMBER "+1555XXXXXXXX"
telnyx-edge secrets add SUPERVISOR_DEVICE "+1555XXXXXXXX"
Use placeholders in source control and store the real values as secrets. The browser softphone also needs a WebRTC login token supplied by the supervisor at runtime.
Production considerations
This sample demonstrates the architecture, but a production coach room should add:
- role-based access for supervisors
- short-lived browser credentials
- policy versioning and approval workflows
- audit retention and deletion rules
- metrics for stream gaps and reconnects
- explicit consent and transcript handling policies
- a post-call source of truth from the Conversations API
The central design still holds: keep observability outside the media path, give each conversation a durable owner, and make human intervention a first-class call participant.
Where to take it next
The same structure can support more than customer service coaching:
- compliance monitoring for regulated calls
- live sales enablement and objection handling
- interpreter or specialist escalation
- healthcare scheduling supervision
- fraud-review assistance
- training dashboards for new agents
You can keep the plumbing and replace the policy inside CoachRoom with the rules that matter to your application.
Top comments (0)