Support teams have a math problem.
Thousands of customer calls happen every week, but supervisors can manually review only a small fraction of them. Even when a call is graded, that score often ends up isolated in a spreadsheet. It cannot tell you whether an agent had one difficult conversation or is developing a pattern.
That creates two distinct engineering problems:
- Coverage: evaluate more completed calls without requiring someone to listen to each recording.
- Memory: preserve each agent's results long enough to identify trends and make better coaching decisions.
The open-source post-call-qa-scoring example addresses both. It uses Telnyx Decision Models to grade transcripts and a durable QAAgent actor to maintain the quality history for each support agent.
The result is not an autonomous disciplinary system. It is a way to surface calls and trends that deserve human attention.
What the application does
When a support call ends, the application:
- receives the transcript from a
call-conversation-endedwebhook - uses
transcription-savedas a fallback when the transcript is not embedded - routes the transcript to a durable actor keyed by the support agent ID
- asks a Telnyx Decision Model for three structured grading signals
- stores the result in private per-actor SQL
- recomputes a five-call rolling average
- flags or clears a coaching recommendation based on that trend
- escalates potential compliance breaches to a manager over SMS
- sends the team lead a daily per-agent digest
Each call ID is stored as a SQL primary key before grading is scheduled. The scheduled task also uses a stable ID, grade:<callId>, so webhook retries do not create duplicate grades.
Why use one actor per support agent?
The central design decision is this:
The actor is the agent's quality profile.
The worker resolves an actor with idFromName(agentId). Every later transcript for that support agent reaches the same durable actor.
const actorId = env.QA_AGENT.idFromName(agentId);
const qaAgent = env.QA_AGENT.get(actorId);
await qaAgent.recordCallEnded(
callId,
transcript,
agentId,
digestEnabled,
);
That actor owns the score history, rolling average, coaching state, breach records, and digest schedule. The application does not have to reconstruct the agent's history from a cache every time another call ends.
Grade the transcript into structured fields
Free-text model output is awkward for QA automation. If the model returns a paragraph, downstream code has to interpret that prose before it can calculate an average or apply an escalation threshold.
The example instead calls the Telnyx Decision Models endpoint:
POST /v2/ai/typesafe/v1/systemone
One request evaluates the transcript against three question types:
-
choice: pass or a specific failure category -
noul: a 0–1 compliance-breach signal -
score: an overall quality score from 0–5
The application can then apply ordinary program logic to typed values:
interface DecisionResult {
choice: string;
noul: number;
score: number;
}
A call can pass its general quality evaluation and still carry a serious compliance signal. The sample therefore treats these fields independently. When noul > 0.8, it records the call in the breaches table and sends an immediate manager-review alert regardless of the pass/fail choice.
Decision Models do not replace the reviewer here. They decide which evidence should be placed in front of one.
Turn individual grades into a trend
After every successful grade, the actor queries its SQL history and recomputes a five-call rolling average.
SELECT ts, choice, noul, score, status, last_error
FROM scores
WHERE status = 'graded'
ORDER BY ts DESC
LIMIT 5;
If the average falls below QA_COACHING_FLOOR, the actor flags the agent for coaching. If later calls bring the average back above the threshold, it clears the flag automatically.
That distinction matters. One low score may be an isolated call. A declining rolling average is a stronger reason to investigate, coach, or review the underlying conversations.
Keep failures visible
The grading task uses bounded retries rather than retrying forever. Transient failures back off for 10, 30, and 60 seconds. If grading still cannot complete, the call is stored with status="ungraded" and a last_error value.
The daily digest can therefore distinguish between a low score and a transcript that was never successfully graded. Failed AI work does not quietly disappear from the reporting pipeline.
Try the complete pipeline without placing a call
The repository includes a synthetic demo endpoint. After deploying the example, call:
curl -X POST https://<edge-function-url>/demo/trigger \
-H "Content-Type: application/json" \
-d '{
"agentId": "demo-agent",
"callId": "demo-001",
"transcript": "Agent: This call may be recorded for quality purposes..."
}'
The endpoint routes the transcript through the same actor, Decision Models request, SQL history, trend calculation, and escalation logic used by the webhook flow. It is intended for demonstration and testing, so you can inspect the architecture before connecting a live call application.
Run the example
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/post-call-qa-scoring
npm install
cp .env.example .env
npm run typecheck
npm test
npm run deploy
For live Decision Models calls and SMS notifications, configure the Telnyx API key, a messaging-enabled Telnyx number, and the team lead's E.164 number as Edge secrets. The repository README documents every supported setting, including the coaching floor and digest hour.
Where to take this pattern next
The sample uses a regional-bank support scenario, but the architecture is useful anywhere conversations need consistent review over time:
- contact-center coaching
- healthcare communication audits
- insurance disclosure checks
- sales-call quality monitoring
- customer-service escalation review
The important pattern is broader than call scoring: produce structured decisions, keep durable history around the entity being measured, and route high-stakes outcomes to a human.
Top comments (0)