Last week I gave a talk on building voice agents, and the part everyone wanted to talk about afterwards was the live demo. Not the slides. The demo.
So this post is that demo, written down. By the end you'll have a voice agent running on your laptop that you can talk to in the browser, interrupt mid-sentence, and watch print its own latency after every turn.
The stack is fully swappable, but here's what I used and why.
- Pipecat as the orchestrator. Open source, Python, and the whole agent reads like a list.
- Smallest AI Pulse for speech to text and Lightning v3.1 Pro for text to speech. These are one of the best models to build voice agents. Both are official Pipecat plugins, and you can swap either one out with a single line.
- Claude Haiku 4.5 as the LLM, because it was the fastest first token I measured from my network on the day. More on that below.
Let's dive in.
A voice agent is not an LLM that talks
Most of us start with this mental model.
One box. Text in, text out. Everything you know about building LLM apps lives in there.
A voice agent looks like this instead.
The LLM is still in the middle, but now it's surrounded by voice activity detection, turn detection, speech to text, text to speech and a real-time transport. And there's a clock. Humans leave about 200 ms between turns on average across ten languages (Stivers et al., PNAS 2009. We won't hit 200 ms with network hops in the loop, but every millisecond past that is felt.
The one idea to take away from this section is that everything hard about voice AI lives in the arrows, not the boxes. Picking good models matters, but the timing between them is the actual engineering.
Why everything has to stream
If each stage waits for the previous one to finish, you add up their full durations.
record → transcribe → think → synthesize → play (slow, every stage waits)
stream → stream → stream → stream → stream (fast, nothing waits for "done")
In a streaming pipeline, speech to text emits partial transcripts while you're still talking, the LLM streams tokens, and text to speech starts speaking on the first chunk of the reply. That last part is why the metric that matters for TTS is time to first byte, not total synthesis time. Once the first audio arrives, the user is listening while the rest is generated.
Pipecat handles all of this for you. Your job is to put the right boxes in the right order.
Setup
You need Python 3.11 or newer, a Smallest AI API key and an Anthropic API key.
mkdir voice-agent && cd voice-agent
python3 -m venv venv
source venv/bin/activate
pip install "pipecat-ai[smallest,anthropic,silero,webrtc,runner]>=1.6.0" python-dotenv
Put your keys in a .env file.
SMALLEST_API_KEY=your_smallest_key
ANTHROPIC_API_KEY=your_anthropic_key
The extras matter. smallest pulls the Pulse and Lightning services, anthropic the LLM, silero the voice activity detector, and webrtc plus runner give you a local browser client and dev server so you don't have to build any frontend.
The whole agent
Here's the full bot.py. It's about 100 lines, and I'll walk through the interesting parts right after.
import os
from dotenv import load_dotenv
from loguru import logger
from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.frames.frames import Frame, LLMRunFrame, MetricsFrame
from pipecat.metrics.metrics import TTFBMetricsData
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import LLMContextAggregatorPair
from pipecat.processors.frame_processor import FrameDirection, FrameProcessor
from pipecat.runner.types import RunnerArguments
from pipecat.runner.utils import create_transport
from pipecat.services.anthropic.llm import AnthropicLLMService
from pipecat.services.smallest.stt import SmallestSTTService
from pipecat.services.smallest.tts import SmallestTTSService, SmallestTTSSettings
from pipecat.transports.base_transport import BaseTransport, TransportParams
load_dotenv(override=True)
SYSTEM_PROMPT = """You are Pixel, a friendly voice assistant.
You are speaking out loud, so follow these rules.
- Never use bullet points, markdown, emoji or headings.
- Keep answers to one or two short sentences.
- Say numbers and dates the way a person would say them.
- If you get interrupted, answer the new question naturally."""
class LatencyMeter(FrameProcessor):
"""Prints time-to-first-byte for every service, every turn."""
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
if isinstance(frame, MetricsFrame):
for d in frame.data:
if isinstance(d, TTFBMetricsData) and d.value > 0:
name = d.processor.split("#")[0].replace("Service", "")
logger.info(f"⏱ {name:<14} time-to-first-byte {d.value * 1000:6.0f} ms")
await self.push_frame(frame, direction)
async def run_bot(transport: BaseTransport, runner_args: RunnerArguments):
stt = SmallestSTTService(api_key=os.getenv("SMALLEST_API_KEY"))
llm = AnthropicLLMService(
api_key=os.getenv("ANTHROPIC_API_KEY"),
settings=AnthropicLLMService.Settings(model="claude-haiku-4-5-20251001"),
)
tts = SmallestTTSService(
api_key=os.getenv("SMALLEST_API_KEY"),
settings=SmallestTTSSettings(voice="meher"),
)
context = LLMContext(
[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": "Say hello in one short sentence."},
]
)
ctx = LLMContextAggregatorPair(context)
pipeline = Pipeline(
[
transport.input(), # mic audio in, VAD runs here
stt, # audio to text, streaming
ctx.user(), # add the user's words to the history
llm, # history to streaming tokens
tts, # tokens to streaming audio
LatencyMeter(), # print the latency budget
transport.output(), # audio out to the speaker
ctx.assistant(), # record what was actually spoken
]
)
task = PipelineTask(pipeline, params=PipelineParams(enable_metrics=True))
@transport.event_handler("on_client_connected")
async def on_client_connected(transport, client):
await task.queue_frames([LLMRunFrame()]) # greet first
@transport.event_handler("on_client_disconnected")
async def on_client_disconnected(transport, client):
await task.cancel()
await PipelineRunner(handle_sigint=runner_args.handle_sigint).run(task)
async def bot(runner_args: RunnerArguments):
transport = await create_transport(
runner_args,
{
"webrtc": lambda: TransportParams(
audio_in_enabled=True,
audio_out_enabled=True,
vad_analyzer=SileroVADAnalyzer(),
),
},
)
await run_bot(transport, runner_args)
if __name__ == "__main__":
from pipecat.runner.run import main
main()
The pipeline is the diagram
Read the Pipeline([...]) list top to bottom. It is literally the diagram from earlier.
Audio comes in from the transport, speech to text turns it into words, the user aggregator adds those words to the conversation, the LLM streams a reply, TTS turns the reply into audio, and the transport plays it. Frames flow down the list, and each processor only cares about the frames it understands.
The line I want you to notice is the last one, ctx.assistant(). It sits after the speaker, not after the LLM. That's deliberate, and it's the reason interruptions work properly. I'll come back to it.
The system prompt is written for the ear
A good text prompt produces bullet points, markdown and long thorough answers. Spoken out loud, that's a 45 second monologue the user interrupts at second four. So the voice prompt forbids formatting, keeps answers short and asks for numbers the way people say them. "$1,234.56" should come out as "twelve hundred thirty four dollars and fifty six cents", not "dollar sign one comma two three four".
VAD lives on the transport
SileroVADAnalyzer() is passed to the transport, not added as a separate step. VAD is a small model that answers one question, "is someone speaking right now?". It gates everything downstream, and it's what lets the agent notice you talking over it.
The latency meter
LatencyMeter is a tiny custom processor. Pipecat already measures time to first byte for every service when you set enable_metrics=True. This processor just catches those metrics as they flow past and prints them. It's the most useful 15 lines in the file, because it turns "it feels slow" into a number per stage.
Run it
python bot.py
Open http://localhost:7860, click Connect and allow the microphone. Pixel greets you.
Wear headphones. I'm serious about this one. Without them, the agent hears its own voice through your laptop mic, thinks someone started talking and stops to listen. To itself.
In production, browser echo cancellation handles this for you through WebRTC. On your laptop during development, headphones are the fix.
Interrupt it
Ask Pixel to tell you a story about a robot who learns to sing. Let it get a few seconds in, then talk over it and ask something else.
It stops immediately and answers the new question. That's barge-in, and two things just happened.
- VAD heard you, and Pipecat cancelled the LLM and TTS work in flight and stopped playback.
- The conversation history only contains the part of the story that was actually played to you.
The second one is the subtle bug in voice agents. If your history says the agent said the full sentence, the agent now believes things the user never heard. Imagine it said "I can book that for Tuesday at 3pm and send a confirmation" and you cut it off at "Tuesday". Your history now claims 3pm was agreed.
That's why ctx.assistant() sits after transport.output(). It records what was spoken, not what was generated.
Read the latency
After every turn the terminal prints something like this. These are real numbers from my laptop in Bengaluru, over WebRTC on localhost.
⏲️ AnthropicLLM time-to-first-byte 874 ms
⏲️ SmallestTTS time-to-first-byte 180 ms
Across a handful of runs on the day, Claude Haiku's first token landed between 0.6 and 1.0 s and Lightning's first audio byte between 150 and 250 ms. That's one laptop and one network, so treat it as an example of what to measure, not a benchmark. When I had a VPN on earlier in the day, the same code was noticeably slower. Your network is part of your latency budget.
A few things I learned from staring at these lines.
- The LLM is usually the biggest slice. On my runs it was several times the TTS number every single turn.
- Measure the model choice, don't assume it. My OpenAI key died on the morning of the talk, so I swapped to Claude. Since I had both, I measured. From my network that day, gpt-4o-mini's first token took 2.3 to 3.0 s and Claude Haiku 4.5 took 1.2 to 1.7 s in the same headless test. Your region may flip that, which is exactly why the meter exists.
- Speech to text doesn't show up as a TTFB line in this setup, because streaming STT reports differently. The meter filters zero values out so you don't read a misleading "0 ms".
If you want a single target, the Voice AI & Voice Agents primer calls 1.5 s voice to voice "an important target to aim for" in its June 2026 edition. Then measure your p95, not your average. The average hides the slow calls, and those are the ones people remember.
Swap a model in one line
Every model in this agent is a single line. To try a different Smallest voice, change the TTS line.
tts = SmallestTTSService(
api_key=os.getenv("SMALLEST_API_KEY"),
settings=SmallestTTSSettings(voice="nolan"),
)
Want a different LLM? Replace the llm = ... line with any other Pipecat LLM service. Different speech to text vendor? Same idea. Restart, reconnect, and notice that you did not touch the pipeline.
That's the whole point of an orchestrator. Models are interchangeable parts. The loop is the engineering you own.
Pipecat or LiveKit?
People always ask this. Both are good, open source orchestrators. Pick one, you never run both. Pipecat is a pipeline of processors you compose in Python, which is why the agent above reads like a list. LiveKit Agents is built around LiveKit's media server and has you define an AgentSession with stt, llm and tts. Smallest has an official plugin for LiveKit too (livekit-plugins-smallestai), so the same model choices carry over.
What's next
You now have a streaming voice agent with barge-in and a latency readout, in about 100 lines. In the next post we'll put this exact agent on a real phone number. The pipeline stays the same and only the transport changes. After that, I'll go through the hard parts of voice agents (endpointing, echo, tool-call latency, 8 kHz phone audio) and how to take all of this to production.
If you build something with this, tag me, I'd love to see it.
References
- Pipecat docs and the Smallest TTS and STT service pages
- Smallest AI docs
- Voice AI & Voice Agents, an illustrated primer
- Stivers et al., Universals and cultural variation in turn-taking in conversation, PNAS 2009
- Full code for this post is in voice-agent




Top comments (1)