DEV Community

Papers Mache
Papers Mache

Posted on

Octree Video Memory Enables Minute-Long Consistency

Many current long‑horizon video generators struggle with growing memory requirements as the number of processed frames increases. GEAR’s octree attention breaks that wall, keeping visual quality high for minute‑scale generations.

Before GEAR, approaches either searched historical context implicitly during denoising or fused all past observations into a persistent global 3D representation. Implicit search suffered from slow memory lookups, while global fusion accumulated geometric drift and could not scale beyond short clips. Those trade‑offs left developers choosing between speed and fidelity.

GEAR produces consistent videos up to 60 seconds long while preserving state‑of‑the‑art visual quality, precise camera control, and revisit consistency on challenging trajectories. It treats per‑frame geometry as a token‑level address, builds an “invisible” octree that accumulates visibility evidence, and routes attention through this constant‑size memory during denoising — thereby avoiding both inefficient searches and error‑prone global fusion[1].

The presented experiments focus on static scene geometry; handling of dynamic objects and their occlusions is not addressed in the paper. The paper does not detail how the invisible octree is constructed, leaving open questions about its integration with end‑to‑end training pipelines. This suggests an open question: can geometry‑as‑address routing be extended to fully dynamic scenes without sacrificing the constant‑memory guarantee?

If GEAR’s paradigm holds, developers should abandon persistent 3D fusion pipelines in favor of token‑level addressable memory for any long‑horizon video synthesis task. Benchmarks that evaluate minute‑long consistency may need to be revisited, and while the authors provide an implementation of the invisible octree, integrating it into other codebases would likely require additional engineering effort.

Will the next generation of video generators discard global 3D memories altogether?

References

  1. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

Top comments (0)