DEV Community

Papers Mache
Papers Mache

Posted on

AI/ML Research Digest — Oct 04, 2026

Cross‑modal token efficiency

Researchers are slashing compute by compressing tokens rather than scaling models. Braco’s four‑step visual coder transforms and orthogonalizes token bases, achieving up to 64× compression with only a 4.8 % drop in accuracy (95.2 % retained) [1].

WUSH‑KV quantizes key–value caches to 2 bits using separate adaptive transforms; the value transform is folded into model weights, cutting memory bandwidth while keeping perplexity competitive [2].

Frozen‑backbone multimodal adapters reuse a static visual encoder and add lightweight adapters for each modality, avoiding expensive fine‑tuning (no citation provided).

MemFold learns a soft, fixed‑budget memory that extends context length without growing compute; downstream reasoning runs up to 4× faster [3].

Modular skill libraries for agents

Skill retrieval pipelines let agents pull reusable procedures on demand. RASO indexes a massive corpus of skills with a bi‑encoder followed by a cross‑encoder reranker, matching the quality of LLM‑generated retrieval while dramatically lowering inference cost [4].

Raven’s host‑agent framework and ActiveSaddler curriculum evolution propose orchestration layers that sequence retrieved skills without hand‑crafted data (citations unavailable).

Skill2Env synthesizes training environments from skill specifications, enabling rapid iteration on new tasks (citation unavailable).

Persistent memory for long horizons

Constant‑size scene memories now support hour‑long video generation and large navigation maps. LOCI builds a hybrid spatial memory that stores compressed scene descriptors; retrieval remains fast regardless of video length [5].

GEAR routes attention through an invisible octree addressed by geometry, keeping per‑frame latents consistent across minutes of video while delivering state‑of‑the‑art visual quality [6].

Lifelong Navigation Memory adds adaptive recall cues to a fixed memory bank, preserving navigation fidelity over extensive routes without linear storage growth [7].

On‑policy distillation for LLM reasoning

Instead of unstable reinforcement fine‑tuning, on‑policy distillation lets a student model learn directly from teacher outputs during inference. Forward KL loss provides stable updates and scales across model sizes [8].

Reverse‑KL regimes transfer knowledge more aggressively when the teacher’s confidence is high [9].

Neighborhood OPSD builds a hierarchical self‑distillation pipeline that improves reasoning accuracy without costly rollout generation (citation does not correspond to the referenced work).

Diffusion advances for language and vision

Diffusion models are being reshaped to match autoregressive quality with far fewer sampling steps. A T5‑Gemma encoder diffusion reduces perplexity on OpenWebText, showing that diffusion can compete in pure language tasks [10].

A one‑step high‑resolution refiner cuts inference latency by 8.9× while preserving image fidelity, proving aggressive step reduction is practical for visual synthesis [11].

Highlighted papers

  • Four‑step visual token coder – compresses tokens 64×, retains 95.2 % accuracy, sets a new efficiency baseline for multimodal models [1].
  • 2‑bit KV cache quantization (WUSH‑KV) – achieves competitive perplexity at 2 bits, folds value transform into weights, dramatically shrinks memory bandwidth [2].
  • One‑step diffusion refiner – reduces latency by 8.9× compared with multi‑step pipelines while keeping image quality high [11].
  • Evidence ledger for multimodal agents – replaces raw dialogue history with a structured evidence log, boosting OmniGAIA performance by 10–15 % without extra token budget [12].
  • Geometry‑latent diffusion (GeoVerse) – injects geometric latent diffusion into video generation, improving PSNR by 2.23 dB and cutting ATE by 32.4 % on challenging benchmarks [13].

Additional notable results

  • Extreme token reduction via self‑distillation – LT‑OPD trains a student with only 5 % of the teacher’s tokens, raising accuracy from 68.6 % to 82.3 % while cutting KV cache use and FLOPs by over 85 % [14].
  • Reusable LLM router (RouteFM) – pretrains a router that generalizes across candidate pools, adding +2.23 quality points on MMR‑Bench and lowering routing overhead versus per‑task training [15].
  • Soft memory for long context (MemFold) – learns a fixed‑budget soft memory via on‑policy optimization, extending usable context length with modest compute and delivering up to 4× speedups in downstream reasoning tasks [3].
  • Octree‑based persistent video memory (GEAR) – treats geometry as an address for attention over per‑frame latents, enabling minute‑long video consistency and state‑of‑the‑art visual quality on long‑horizon benchmarks [6].
  • Cost‑effective skill retrieval (RASO) – bi‑encoder plus cross‑encoder pipeline retrieves relevant procedural skills from a massive corpus, matching LLM‑mediated retrieval performance while substantially reducing inference cost [4].

References

  1. Beyond Selection: Token Parameterization for Extreme Visual Token Compression
  2. WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
  3. MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
  4. Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
  5. LOCI: Spatial Linear Memory for Streaming World Models
  6. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
  7. NavHarness: Towards Lifelong Embodied Navigation
  8. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
  9. Scaling Properties of Same-Family On-Policy Distillation
  10. Scaling and Distilling Text Embeddings for Better Diffusibility
  11. SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
  12. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
  13. GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
  14. Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
  15. Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Top comments (0)