I run OpenClaw as my main agent harness, with the voice layer on Discord. When I first turned voice on, I said hello and waited about eight seconds for a reply. Eight seconds is long enough to assume it doesn't work.
It works. The delay was two configuration defaults. Neither was the model.
OpenClaw is the open-source agent harness I run my workflow on; the voice layer sits on Discord. The config keys are OpenClaw's, but the cause — turn-taking defaults — is generic.
Two defaults caused most of the delay
A fixed silence wait. captureSilenceGraceMs defaults to 2000 ms. It waits two full seconds after you stop talking before it decides your turn is over. Normal conversation leaves a few hundred milliseconds. Two seconds reads as a pause.
A full-agent round-trip on every turn. The realtime voice model consulted the entire agent before it spoke. I'd say "hello"; OpenClaw would run the whole agent — tools, memory, the lot — and only then would the voice layer talk. That round-trip is the larger half of the delay — the same failure I wrote about from the theory side: the pause lives in the turn-taking, not the LLM.
Both are turn-taking settings, not model quality.
The config
Everything sits under channels.discord.voice in ~/.openclaw/openclaw.json. Before → after:
| Setting | Before (default) | After | Why |
|---|---|---|---|
captureSilenceGraceMs |
2000 |
800 |
Wait a beat, not two seconds |
realtime.consultPolicy |
always |
auto |
Voice layer handles the conversational filler; call the agent only when the turn needs it |
realtime.bargeIn |
(on) | true |
Pinned — cut it off mid-sentence |
realtime.minBargeInAudioEndMs |
250 |
0 |
Interrupt with no 250ms tail gate |
realtime.instructions |
ends …conversational. Never open with filler.
|
ends …conversational.
|
Lets it use a short backchannel ("one sec") while an agent consult runs |
model (agent behind voice) |
openai/gpt-5.6-terra |
deepseek/deepseek-v4-flash |
Cheaper, faster reasoning behind the realtime front end |
The one that mattered most was consultPolicy: "auto". The realtime model (gpt-realtime-2.1) is only the voice layer — turn-taking, barge-in, playback. The reasoning comes from the agent, which the realtime model calls as a tool. On "always", every turn called the agent before producing audio. On "auto", the voice layer handles the conversational filler itself and calls the agent only when the turn needs it.
The block, after:
"channels": {
"discord": {
"voice": {
"enabled": true,
"mode": "agent-proxy",
"captureSilenceGraceMs": 800,
"realtime": {
"provider": "openai",
"model": "gpt-realtime-2.1",
"speakerVoice": "cedar",
"providers": {
"openai": { "apiKey": "***}" }
},
"instructions": "Always speak English unless the user is clearly speaking another language to you. Keep spoken replies short, natural, and conversational.",
"consultPolicy": "auto",
"bargeIn": true,
"minBargeInAudioEndMs": 0
},
"model": "deepseek/deepseek-v4-flash"
}
}
}
You need an API key for the realtime provider. OpenAI here; Grok if you want to price that side.
Result
The realtime model answers first, then reports what it called. The pause is gone and the back-and-forth is responsive.
Cost
One session, over the drive from home to the office: $3.25, mostly input — the base listening. Roughly thirty cents a minute on the realtime model while you talk.
Known issues
- Over Bluetooth, reply audio cuts out in a way music doesn't. The link is less stable than it should be.
- Long dictation and quick back-and-forth are different modes. This setup is good at the back-and-forth; for long dictation I use a separate dictation tool.
If voice on your agent feels slow, check the silence grace and whether the voice layer consults the full agent every turn before you blame the model. If you've hit a different turn-taking bottleneck, I'd like to know which one.
- My default model stack for AI agent work — why a cheap, fast model behind the realtime front end holds up.
- How much it costs to run a capable AI agent each month — the fuller version of the $3.25-a-session cost.
- What is an AI work harness — what the "agent behind the voice" actually is.
Originally published at engineering.kenmazaika.com.
Top comments (0)