Multimodal bias and low‑disturbance control
Omni‑modal language models consistently favor visual input over audio (a “perceptual” bias) and favor propositional audio content (an “evidence‑form” bias). The bias appears already in the first transformer layers, so it can be read out with little computation [1]. This early emergence means we can intervene without harming downstream tasks.
Fisher‑Rao geometric analysis shows that next‑token distributions of different model families share essentially the same metric [2]. Because the geometry is shared, a single low‑disturbance control method works across architectures: it nudges a token distribution while leaving most of the representation untouched. The result is a principled way to correct bias or enforce safety constraints with minimal performance loss.
Scalable autonomous‑agent architectures
Large perception datasets and memory mechanisms are now the main drivers of progress in embodied AI.
- The WROP object‑permanence dataset, paired with a 16 B parameter video world model, achieves the highest scores on an exhaustive reasoning exam [3]. This shows that grounding visual streams in physical‑object concepts is becoming tractable at scale.
- Adding a fixed‑size episodic buffer to visual‑language agents boosts success on VLA tasks by more than sevenfold compared with stateless baselines [4]. Short‑term storage therefore appears essential for coherent, multi‑step interaction.
- Agent‑Editing introduces a “state revision” operation that lets an agent modify its internal snapshot before continuing execution; this yields measurable gains across several benchmark suites [5]. It demonstrates that agents can benefit from explicit self‑correction mechanisms.
- Realtime‑Venus implements a full‑duplex audio‑visual dialogue system that perceives continuously and delegates tool use asynchronously. It sets new state‑of‑the‑art results on multiple video and speech benchmarks [6], confirming that tightly coupled perception–action loops are viable at real‑time speeds.
Formal specification for skill verification
Instead of checking raw execution traces, recent work validates reusable skills with formal specifications.
- SkillSpec treats each autonomous‑agent skill as a Hoare triple (precondition → action → postcondition). It automatically derives “intent” masks that filter out contextual bias and “FactSpecs” that capture expected world changes. On a real‑world suite of 515 skills, the system reaches 61.2 % precision in matching intended outcomes [7].
- Code2Skill mines over one million atomic operations from open‑source codebases, assembling them into a CodeSkillBank. When agents draw skills from this bank, average benchmark performance improves by 11.7 % across heterogeneous tasks [8]. The result is a scalable pipeline for turning code fragments into verified action primitives.
Alignment and reward modeling
RewardVerse applies rubric‑guided scoring to video‑based reward models, dramatically reducing scalar drift that typically plagues long‑horizon learning. On the EvalVerse alignment suite it achieves state‑of‑the‑art results [9], suggesting that structured evaluation can keep learned rewards stable over time.
Why these matters:
Detectable multimodal bias lets us correct models before they affect downstream behavior, and Fisher‑Rao control gives a mathematically sound tool for doing so. Scalable agents with memory and self‑editing are the first steps toward robust embodied AI that can plan, remember, and recover from mistakes. Formal skill specifications turn ad‑hoc code snippets into reliable building blocks, accelerating development cycles. Finally, stable reward modeling is crucial for aligning long‑running agents with human intent. Together, these advances tighten the feedback loop between model capabilities and trustworthy deployment.
References
- Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
- The information geometry of large language models is shared, learned, and controllable
- Training Object Permanence in World Models
- MemBodied: Recurrent Associative Memory for Vision-Language-Action Models
- Agent-Editing World Model: Rethinking World Modeling for LLM Agents
- Realtime-Venus: A full-duplex interaction system with asynchronous delegation
- SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
- Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
- RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
Top comments (0)