DEV Community

Papers Mache
Papers Mache

Posted on

5% Tokens, 14% Accuracy Gain via Self-Distillation

LT‑OPD keeps only five percent of visual tokens yet lifts accuracy from 68.6 % to 82.3 %. The result flips the conventional wisdom that aggressive token pruning inevitably harms performance, and it does so while slashing memory traffic and compute overhead.

Previous visual-token reduction approaches often suffered notable accuracy degradation when token budgets were reduced aggressively, motivating the need for more effective strategies such as LT‑OPD. The community therefore assumed that extreme budgets were impractical for high‑quality multimodal reasoning.

Across nine benchmarks on Qwen3.5-4B, LT‑OPD raises average retained performance under 5 % visual-token retention from 68.6 % to 82.3 %, outperforming training‑free, training‑based, and reinforcement‑learning baselines at the same budget. The student model rolls out responses using only a tiny slice of visual evidence while a frozen full‑token copy supplies distributional supervision along those trajectories. This on‑policy self‑distillation recovers most of the capability lost to token compression.

The same framework trims KV‑cache usage by 85.2 % relative to the full‑token model. Because cache size grows linearly with token count, cutting tokens shrinks memory pressure dramatically. Practitioners can therefore fit longer context windows or larger batch sizes on unchanged hardware.

Prefill FLOPs drop by 85.4 % without adding inference latency. Fewer tokens mean a shorter initial matrix multiplication chain, yet the generation phase runs at the same speed because the student already learned to compensate for missing visual cues. The net effect is a faster, cheaper forward pass that does not sacrifice answer quality.

The gains hinge on a budget‑level curriculum that gradually shrinks the token allowance during training. “To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training.” [1] This adds a non‑trivial scheduling component and requires a frozen full‑token teacher, leaving open how well the approach scales to other modalities or higher token budgets.

If extreme pruning can simultaneously accelerate inference and improve accuracy, benchmark suites should include a five‑percent token track as a first‑class setting. Model developers can adopt LT‑OPD as the default visual front‑end, turning what used to be a latency‑accuracy trade‑off into an outright win.

References

  1. Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

Top comments (0)