DEV Community

Daniel Varela
Daniel Varela

Posted on

Three real sessions on a CPU voice agent API

Three recorded sessions against the Lokutor voice agent API. Speech-to-text, LLM, and voice all run on CPU. The millisecond figure on screen is measured server-side: first audio after end of turn detected.

1. Booking a dentist, in Spanish

Two agents on the same API: a caller and a clinic receptionist. The caller asks for a cleaning this week, gets offered Thursday 5:30 PM or Friday 10 AM, picks Thursday, asks the price (45 euros), and books. Latency shown per turn: 207 ms, 206 ms, 144 ms. 40 seconds.

2. One voice, four languages

Same agent, same voice, switching live between Spanish, Catalan, Galician, and Basque. Latency on screen: 204 ms in Spanish, 144 ms in Catalan, 67 ms in Galician. The API also handles English, Portuguese, French, Italian, and German. 31 seconds.

3. Two agents debating

Two Lokutor agents debate in Spanish whether AI will end humanity. Unscripted: each one hears the other and answers on its own. Humanity survives, for now. 36 seconds.

Run it yourself

Free API key at app.lokutor.com, docs at docs.lokutor.com.

Top comments (0)