A voice agent can answer correctly and still make a conversation frustrating. It can wait too long before replying, keep speaking after an interruption, or confirm the wrong appointment time after a correction.
Those failures need more than a transcript review.
Why build a tool for this? I'm building a voice agent as a hobby project: real-time STT, LLM and TTS, the whole loop. Open-source voice evals do exist. Pipecat ships evals that script turns against its own transport, and ServiceNow's EVA is an enterprise benchmark suite where the agent under test runs inside EVA's harness. Neither scores transcription accuracy against ground truth (EVA keeps WER diagnostic-only; Pipecat's judge is told to tolerate ASR errors), neither surfaces per-stage latency, and both bind you to their world. I wanted the opposite trade-off: deterministic scoring (jiwer WER, exact outcome matching, no LLM judge), per-stage latency budgets, barge-in stop time, CI gates, and a probe that talks to any WebSocket agent through a declarative protocol map. So I built voice-evals to turn recorded calls and scripted live conversations into repeatable evaluations.
The project is a Python CLI and library, released under Apache 2.0. Its two entry points use the same evaluator: replay a saved call, or place a scripted call through a compatible WebSocket transport and export the recording for later replay.
Start with a recording
The fastest way to try it needs no API keys. Clone the repository for the example data, then install the package:
git clone https://github.com/Gjusev/voice-evals.git
cd voice-evals
python -m pip install voice-evals
voice-eval run evals/data/demo_calls.jsonl
The three-call synthetic demo includes these results:
samples=3 failures=0
wer mean=0.0303 max=0.0909
task_completion=0.6667
fact_coverage=0.8333
hallucination_rate=0.3333
e2e_ms p50=950.0 p95=1103.0 p99=1116.6
These values exercise the evaluator. They are not a ranking of speech providers, and failures=0 does not mean every task succeeded: completion is reported separately.
Each JSONL row combines a scenario, reference and observed transcripts, an outcome, and optional timings and interruption observations. When your data is ready, define the thresholds that matter to your application:
voice-eval run calls.jsonl --max-wer 0.05 --min-task-completion 0.90 --max-e2e-p95-ms 1000 --json --output result.json
A breached gate returns exit code 1. The numbers above are example thresholds, not universal voice UX targets.
Exercise the conversation
Replay is useful for regression testing. To observe an agent's behavior during a conversation, the probe acts as a scripted caller.
The bundled appointment scenario includes a conditional branch and a caller interrupting to correct the requested time. Try its offline mock first:
python -m pip install "voice-evals[probe]"
voice-eval probe src/voice_evals/resources/scenarios/appointment-v2.json --mock --output-dir out/probe-demo --max-wer 0.05 --max-barge-in-stop-ms 500
The mock run takes about 25 seconds. It tests the harness and simulates timing; it does not make a real provider call.
For a real probe, configure the caller provider and a WebSocket endpoint that matches a protocol map. The transport currently supports mono s16le PCM at 16 or 24 kHz. The repository includes a local reference agent so you can exercise the socket path before integrating your own service.
Caller options include ElevenLabs, prepared audio fixtures and an HTTP TTS service. Open-model examples show how to connect separately hosted models; the evaluator does not download and serve those models itself.
Keep the evidence
A probe writes a recording directory:
manifest.json configuration, provenance and hashes
events.jsonl normalized event journal
calls.jsonl exported replay records
result.json scores and probe observations
audio/caller/ sent audio
audio/agent/ received audio
For an eligible, scored session, the exported calls replay to the same legacy scores:
voice-eval run out/probe-demo/calls.jsonl
This is reproducibility of a recording's evaluation. A new live call may behave differently because the agent, speech services and network can vary.
A completed session can fail its gate
The repository records a real run through ElevenLabs TTS, a WSS connection and a reference agent using Scribe STT. The session completed. It produced a WER of 0.1379 and an observed barge-in stop time of about 232.6 ms.
Its WER threshold was 0.10. The gate failed.
That failure is useful evidence. The recording survived, the score could be replayed, and the threshold breach stayed visible. This was one session with a scripted reference dialogue policy, so it establishes a tested transport path, not a claim about general agent intelligence.
One more detail about aggregation is worth spelling out. Probe turn-level latency percentiles describe individual observed turns. Legacy replay percentiles describe the per-call timing values exported to the dataset. They answer different questions and should not be substituted for one another.
Be precise about what is measured
Four boundaries shape the output:
- Hidden LLM and TTS stages are not reconstructed from client events. Unobserved stage timings stay unavailable.
- If an interruption never confirms that audio stopped, the result is censored with a lower bound. It is not assigned a fabricated duration.
- Task completion compares recorded and expected outcome labels. Required and forbidden facts use case-insensitive substring checks. The field called
hallucination_ratemeasures configured forbidden phrases, not open-ended factual accuracy. - Sessions that lack the evidence required for scoring report
NOT SCOREDand list the reasons.
The frozen regression corpus contains 123 synthetic calls, including hand-checked WER examples and controlled fault cases. It validates the harness rather than benchmarking live models.
Try it
Start with the repository and quick start, or inspect the Kaggle offline kernel source. The offline Kaggle notebook runs with internet disabled from a checksummed wheel bundle and passed its full check suite on 2026-10-04; the live-probe notebook falls back to a clearly labeled mock run when secrets are absent.
I'd especially like feedback on protocol adapters and failure scenarios: interruptions, corrections, missing facts, and the cases that made your own voice agent fail in practice.
Top comments (0)