DEV Community

Cover image for qwen-audio-agent keeps the voice talking while your coding agent works
Reno Lu
Reno Lu

Posted on

qwen-audio-agent keeps the voice talking while your coding agent works

Reading the qwen-audio-agent README, what stands out is the architecture: a realtime voice layer sits in front of an agent you may already run, such as Qwen Code, OpenCode, Codex or Claude Code, keeps the spoken conversation going, and hands tool work to that backend. The README frames the goal as presence: conversation should not grind to a halt because the Agent is looking something up, calling a tool, or working on a task.

Two lanes, one assistant

The architecture section describes the routing. Questions that can be answered directly are answered immediately; when tools or sustained processing are needed, the task is delegated to the backend Agent. The README adds that throughout, the user always faces the same assistant.

The feature list fills in the concurrency. Frontend conversation and background tasks run in parallel, and you can ask about progress or cancel at any time. You can create multiple independent tasks that the backend Agent executes asynchronously, with continuous status tracking. When a task completes, its result returns to the current conversation, where you can ask follow-up questions or request modifications. The voice side is described as full-duplex realtime interaction with natural interruption and sustained multi-turn conversation.

The README calls this a "foreground conversation + background task" design.

Bring your own backend

The backend side is where I would spend evaluation time. The README says integration reuses your preferred Agent's model configuration, tools, MCP, Skills, and authentication. Backends connect over ACP, and the support table lists how each one gets there: native ACP for Qwen Code, OpenCode, Qoder, MiniMax Code, Kimi Code, Hermes, CodeBuddy and DeepSeek; a built-in ACP bridge for OpenClaw; and an external ACP adapter for Codex, Claude Code and Pi.

Each row carries a rating. According to the README, five stars mean a thoroughly tested recommended integration, while four stars indicate active development or not yet fully verified. Codex, Claude Code, Pi, DeepSeek, MiniMax Code, Hermes and CodeBuddy sit at four stars today, so if your team works in one of those, read the backend documentation before planning around it. There is also a "None" row for frontend-only mode, which needs no config.

The News section records that backend Agents were unified under the ACP architecture in v0.9.0, when the project was open-sourced. The v2.0.0 entry says the next major version is under active development, with ongoing work on the Agent architecture, task lifecycle, multimodal input, memory, and extensibility.

Getting it running

The README asks for Node.js 22.22.2+ or 24.15.0+ with npm 10+, and recommends the global npm install:

npm install -g qwen-audio-agent
Enter fullscreen mode Exit fullscreen mode

Then qwenaudio config creates your config. The example sets DASHSCOPE_API_KEY, QWEN_AUDIO_REALTIME_MODEL for the voice frontend, AGENT_PROTOCOL for the backend (leave it empty or set it to none for frontend-only mode), and an optional QWEN_AUDIO_AGENT_BACKEND_MODEL. The comment on that last variable says explicit values use standard ACP, while an empty value reuses the Agent config.

You start the Gateway in one terminal with qwenaudio and the TUI in another with qwenaudio tui, or use qwenaudio webui for the browser UI. DashScope realtime voice is the default. The README lists Speech-to-Speech and ModelBest MiniCPM-o 4.5 as alternatives, with local or hosted endpoints selected through their service URL, and the v1.3.0 note describes the speech-to-speech integration as supporting fully local VAD, STT, LLM, and TTS.

A desktop app is also available. It provides a persistent floating voice orb with a built-in Gateway, automatic idle sleep, local voice wake, and customizable appearance, and the README includes build scripts for macOS, Windows and Linux.

Beyond the desktop

The README says the current framework focuses on desktop productivity, then argues the same split can expand to other scenarios. Its table marks a smart cockpit example and the VoiceMem and LightRAG integrations as available, an AI Passport hardware card as experimental and currently half-duplex only, customer support and embodied intelligence as planned, and a livestream assistant as exploratory.

The smart-cockpit reference scenario ships in the repository, and the README describes its cockpit UI, small A2A Agent, and cockpit service as customer-replaceable examples. LightRAG connects through the generic KnowledgeProvider, and the README notes that the core npm package includes neither LightRAG nor VoiceMem Python code or dependencies. Lightweight Markdown memory remains the default.


GitHub: https://github.com/QwenAudio/qwen-audio-agent


Curated by Agent Palisade — practical AI for small and mid-sized businesses.

Top comments (0)