On Sunday I built Vaani, a phone receptionist demo for a dental clinic in Surat that doesn't exist. It has a keypad menu for Gujarati, Hindi or English, tools to book, move and cancel appointments in a diary, and a route to the clinic's own front-desk number when a person should take over. The diary lives in memory and resets on restart, and there's no live number yet. So everything below is how it's built and what the tests check, not what callers did.
A lot of the design is about what happens while the model is still working. A turn can include tool calls (look up Tuesday, check which dentist is free, book the slot), and a phone line has no typing indicator to cover the wait.
A turn that outlives the webhook
With Twilio's <Gather> speech input, a call is a series of HTTP requests. Twilio posts what the caller said, you answer with TwiML (say this, listen again), and Twilio allows a request up to 15 seconds.
The simple version holds that request open until the model has finished its whole reply. In Vaani a turn isn't tied to one request. It's an object that lives in memory across several:
interface PendingTurn {
promise: Promise<TurnResult | null>;
startedAt: number;
/** Webhooks answered so far for this turn. */
waits: number;
/** Every sentence the agent has produced this turn, in order. */
sentences: string[];
/** Server-side cursor into `sentences`. */
spoken: number;
// (plus a wake-up callback, trimmed here)
}
When the transcription comes in, the turn starts and the first request gives the model up to 900 ms. If the turn is done by then, its reply goes back. If not, the response carries the complete sentences the agent has written so far, or a short holding line if there aren't any, and ends with a redirect to the server. The next request picks up the same turn and waits again, up to 5 seconds this time, returning sooner if the turn finishes or a new sentence turns up. Round it goes. That spoken counter is bookkeeping on the server, not proof of what a caller heard.
Those are server-side waits, not the moment a caller hears something. Preparing audio can add up to another 1.5 seconds (more on that below), and Twilio still has to fetch and play the response. Measuring what a caller actually hears is the next job.
The first holding line changes from turn to turn ("One moment.", "Let me check.", "Okay, checking."). If a turn needs a second wait with nothing new to say, it's "Still on it, thanks for waiting."
There's a ceiling. The turn has a 20 second budget, checked each time a request picks it up. Run past it, or throw an error, and the call goes to the handoff path, which dials the clinic's number. If there's no number configured or nobody picks up, it notes a callback instead.
Sentences out of a token stream
The best thing to say during the first wait is the agent's own first sentence, if it has written one. A reply that starts "Let me check Tuesday for you" is a perfectly good thing to hear while the tool runs. That means cutting the token stream into sentences as it arrives, in three scripts.
A full stop is the awkward one. It turns up in "499.00", in "Dr. Shah", and at the end of a half-streamed "4." that's about to become "4.30". So a full stop only counts once whitespace has arrived after it, which handles the numbers. Titles need their own check, because "Dr." already has a space after it, so there's an abbreviation list with the English, Gujarati and Hindi ones ("Dr.", "ડૉ.", "डॉ." and friends).
Hindi also ends sentences with the danda, "।", and it can arrive with no space after it. So the whitespace rule is for the full stop only. A danda doesn't need a space after it, though it still waits for a little more input (so trailing punctuation or a closing quote gets absorbed) and for the length check below to pass:
const TERMINATORS = new Set([".", "!", "?", "।", "…"]);
// ...
if (ch === "." && !/\s/.test(next)) continue;
if (ch === "." && isAbbreviation(this.buffer.slice(0, i + 1))) continue;
if (end + 1 < MIN_SENTENCE_CHARS) continue;
Two limits sit around that, both counted in JavaScript string length rather than visible letters. A boundary less than 12 units in is skipped, so tiny fragments don't go out on their own. If the buffer passes 240 with no ending in sight, it's cut at the last space at or before 240, or hard at 240 if there's no usable space, so audio keeps moving. When the model finishes, whatever's left is flushed, short or not.
The voice was the slow part in Gujarati
Voices are picked by language. With an ElevenLabs key set, English and Hindi use their Flash v2.5 model, which my build notes put at about 0.6 to 0.8 seconds a line. For Gujarati I tried eleven_v3, and in that one session it took about 8 seconds a sentence. That's far too slow for a call, so Gujarati goes through Twilio's own Google Chirp 3 HD voice instead, which Twilio synthesises itself.
For English and Hindi, each sentence is queued for synthesis as soon as it's written (two requests to the provider at a time, so one can wait its turn). Before a response goes back, the router waits up to 1.5 seconds for that audio. If a clip is ready it's played. If not, the line is spoken with Twilio's Chirp voice. The fixed prompts (the menu, consent, the holding lines) are queued for synthesis at startup and cached, with the same fallback if a clip isn't ready.
What's tested and what isn't
There are 45 tests, covering things like sentence boundaries, the diary and cancellation rules, and which fields are allowed into diagnostics. A separate eval script has 11 scripted conversations across the three languages, checking which tools the agent called and what state the diary ended up in. They're text conversations, not recorded phone calls.
It's still turn based. Replies from a finished turn go inside the speech gather in the TwiML, which is what lets a caller talk over them. The lines sent during an unfinished turn sit outside it, before the redirect. The tests check the TwiML that's generated; how interruptions behave on a real call is still to be tried. Full duplex audio would mean Twilio Media Streams and a streaming speech stack, and the agent sits behind an interface so that can be swapped in later.
Next is a real number, and timing what a caller actually hears on an Indian mobile network.
Top comments (0)