Visual bias now outweighs any other modality signal in the majority of omni‑modal language models, yet the field still lacks a proven way to excise it from the earliest representation layers. The new Tri‑PvP benchmark reveals that early‑layer preferences are both measurable and only partially suppressible, forcing a rethink of how we certify multimodal neutrality.
Previously, researchers evaluated modality bias on mixed benchmarks that conflate perceptual cues (raw images or audio) with propositional statements (“this is a dog”), making it impossible to attribute errors to the visual stream itself. Standard mitigation attempts have focused on post‑hoc output filtering or fine‑tuning on balanced data, but they never demonstrated that bias disappears before the final classifier.
Visual modality dominates the answer selection in omni‑modal LLMs, often exceeding a 60 % share of the bias signal. “Among the 20 (model evidence-type) bars in Figure 3, BIAS_IMAGE is the dominant bias label in 18… The magnitude of BIAS_IMAGE frequently exceeds 60 %” shows that images drive the model’s final answer far more than audio or text cues [1].
Modality‑bias signals are already linearly recoverable from the first few transformer layers. “Modality‑bias information generally becomes linearly decodable from early layers rather than only at the output stage” confirms that a simple linear probe can expose bias without waiting for the model’s decision head [1].
A contrastive decoding based low‑disturbance adjustment reduces bias while preserving task performance, dropping OmniBench accuracy by only one point (38.4 % → 37.4 %). The authors report that “contrastive decoding mitigates this bias at inference without parameter updates while keeping OmniBench performance (37.4 % vs 38.4%)”, demonstrating a mathematically grounded intervention that barely nudges downstream metrics [1].
The remedy does not fully eradicate early‑layer bias; residual visual preference remains detectable by the same linear probes, and the contrastive decoding adjustment only softens rather than eliminates it. This suggests that future work must combine representation‑level regularization with architecture‑aware training objectives if complete neutrality is required.
Tri‑PvP should become a mandatory sanity check before deploying omni‑modal models, and developers can integrate early‑layer probing plus contrastive decoding into their pipelines to certify that visual bias stays below the 60 % dominance threshold while keeping downstream scores essentially intact.
Top comments (0)