We evaluate four LVLM backbones (LLaVA-1.5-7B, LLaVA-1.5-13B, Qwen-2.5-VL, InternVL-2.5) on CHAIR, POPE and MME, using greedy decoding. Following the analysis in the previous section, our intervention replaces the value vectors of the bottom-25% low-attention tokens during the focus phase with an image-level mean value vector, computed from a single forward pass. No training, gradient updates, or iterative optimization is required.
Focus Matters: Attention-Value Dynamics for Hallucination Mitigation in Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) have achieved impressive progress in multimodal reasoning, yet they remain prone to object hallucinations, generating descriptions of objects that are not present in the input image. In this work, we investigate hallucination from the perspective of attention-value dynamics inside LVLM vision encoders. We identify a consistent three-phase structure of visual processing—diffusion, focus, and rediffusion—and show that the focus phase is where attention most clearly separates strongly and weakly supported visual tokens. However, low attention does not necessarily imply negligible downstream influence: low-attention tokens in the focus phase can still exert non-negligible value-side influence on the attention output relative to their small attention mass. Through controlled phase-wise value interventions, we find that hallucination behavior is particularly sensitive to the value content of low-attention tokens during this phase. Replacing or neutralizing these values reduces hallucination metrics while largely preserving grounded object evidence. A token-level teacher-forcing analysis further shows that the intervention reduces the probability of hallucinated object tokens with much smaller effects on ground-truth object tokens. In addition, Visual Attention Ratio (VAR) analysis shows that focus-phase intervention is accompanied by increased attention to visual tokens during decoding. Based on these observations, we instantiate a simple training-free inference-time intervention that replaces focus-phase low-attention values with an image-level mean value vector using statistics from a single forward pass. Experiments across multiple LVLM backbones demonstrate that this analysis-derived intervention reduces object hallucination with negligible additional runtime and remains compatible with existing inference-time mitigation methods.
Analysis
01 · Three-Phase Structure
Three-Phase Attention Structure in Vision Encoders
We analyze how attention distributions evolve across vision-encoder layers and find a consistent depth-wise pattern across multiple LVLM backbones, independent of architecture or scale.
We summarize layer-wise attention concentration as \(R^{(\ell)} = \mathbb{E}_h[M^{(\ell,h)}] / \mathbb{E}_h[H^{(\ell,h)}]\), the ratio of the maximum attention score to attention entropy. Tracking \(R^{(\ell)}\) across layers reveals three distinct phases:
- Phase 1 · Diffusion: in early layers, attention is broadly distributed across visual tokens.
- Phase 2 · Focus: in intermediate layers, attention becomes sharply concentrated on a small subset of tokens, most clearly separating strongly and weakly supported tokens.
- Phase 3 · Rediffusion: in later layers, the concentrated pattern spreads out again.
Quantitative Results
Hallucination Mitigation & Inference Efficiency
We compare against the original models and AUE, which estimates unreliable visual tokens via iterative PGD-based optimization.
- Consistent hallucination reduction: lowers \(\text{CHAIR}_S\) and \(\text{CHAIR}_I\) across all four backbones, comparable to AUE while generally preserving caption fidelity (F1) better.
- Preserved visual recognition: maintains or improves F1 and POPE accuracy across Random, Popular, and Adversarial splits, indicating that grounded visual evidence is retained.
- Negligible runtime overhead: on Qwen-2.5-VL, AUE inflates inference from 2.60s to 33.46s (~12.9× slower); ours runs in 2.71s. On InternVL-2.5, AUE takes 41.29s vs. our 2.44s (~16.9× slower).
General Capability on MME
Beyond hallucination-specific benchmarks, we check on MME whether the intervention disturbs broader multimodal capabilities. We compare per-category MME scores on Qwen-2.5-VL and InternVL-2.5.
- Profile largely preserved: the radar shapes follow the original models across both perception-oriented categories (existence, count, position, color, OCR, scene, etc.) and cognition-oriented categories (commonsense reasoning, numerical calculation, text translation, code reasoning).
- No global suppression of visual signals: hallucination reduction on CHAIR and POPE is not achieved at the cost of degraded general multimodal performance.
Qualitative Results
BibTeX