Realtime‑Venus sets new top scores on several video understanding and spoken dialogue benchmarks. By coupling a 9 B audio‑visual model with a dedicated speech model in a full‑duplex, continuously perceived architecture, it delivers native speech generation while delegating tool use asynchronously.
Before this work, most multimodal agents operated in half‑duplex mode or relied on post‑hoc reasoning pipelines that paused the conversation for background computation. Systems such as Gemini 3.1 Live and GPT‑4o demonstrated impressive single‑modal performance but fell short of seamless, real‑time interaction across visual and auditory streams.
Realtime‑Venus‑Omni tops six of eight video benchmarks, reaching 70.2% on StreamingBench. “Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%).” [1] The model’s gains translate into consistent leads over prior online baselines across diverse streaming scenarios.
Realtime‑Venus‑Audio leads across four speech tasks, scoring 78.0% on MMAU. “Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81.” [1] Its performance closes the gap between perception‑only models and fully conversational agents that must also generate speech.
On Full‑Duplex‑Bench v1.5, Realtime‑Venus‑Audio answers 75% of user interruptions and sustains continuation rates up to 97%. “On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.” [1] This demonstrates that high‑quality dialogue can persist even when users interject or speak over the system.
The paper evaluates only 9 B checkpoints in controlled benchmark environments, leaving open questions about scaling to larger models, latency under varied network conditions, and robustness to out‑of‑distribution visual scenes. This suggests that future work should probe how asynchronous tool delegation behaves when external APIs exhibit unpredictable delays, and whether similar architectures can retain their advantages at the hundred‑billion‑parameter scale.
If these results hold in production settings, benchmark suites for multimodal agents must incorporate full‑duplex interruption metrics as a first‑class evaluation criterion; otherwise progress will be measured on incomplete proxies that ignore real‑time conversational continuity.
Top comments (0)