DEV Community

Papers Mache
Papers Mache

Posted on

Episodic memory boosts visual agents sevenfold

Adding a fixed‑size episodic memory buffer lifts visual‑language‑action success rates by more than sevenfold — MemBodied demonstrates that a tiny cache can dominate stateless baselines [1].

Before this work, most vision‑language‑action (VLA) policies treated each observation in isolation or relied on naïve recurrent hidden states; the standard π₀ stateless controller and a vanilla recurrent memory were the de‑facto references for history‑dependent manipulation.

MemBodied reaches 50.0 % mean task success, exceeding the stateless policy by 43.6 percentage points, which translates to roughly a 7.8× improvement over the 6.4 % baseline [1].

On three real‑robot tasks the same model jumps from 3.33 % to 26.67 % mean success, an eightfold lift that persists outside simulation [1].

It also cuts inference latency by 91.9 % and reduces peak GPU memory usage by 9.5 % relative to a native memory‑augmented alternative, proving the buffer’s efficiency as well as its efficacy [1].

The paper does not explore how performance scales when episodes stretch far beyond the fixed capacity, nor does it evaluate tasks that require hierarchical or long‑range reasoning; the reported gains are confined to five RMBench scenarios and a handful of physical trials, leaving open whether larger or adaptive buffers would preserve the same ratio.

Practitioners should replace stateless VLA controllers with a lightweight episodic buffer as the new default and re‑run established benchmarks, because the assumption that recurrent hidden states alone suffice for history‑dependent tasks is now empirically challenged.

Will the next generation of household robots rely on a few hundred kilobytes of episodic cache rather than gigabytes of rolling context?

References

  1. MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

Top comments (0)