Focus Matters: Attention-Value Dynamics for Hallucination Mitigation in Vision-Language Models

Sohyeon Kim1 Sang Yeon Yoon2 Kyeongbo Kong1†

1Pusan National University 2Pukyong National University

†Corresponding author

NeurIPS 2026

Abstract


Large Vision-Language Models (LVLMs) have achieved impressive progress in multimodal reasoning, yet they remain prone to object hallucinations, generating descriptions of objects that are not present in the input image. In this work, we investigate hallucination from the perspective of attention-value dynamics inside LVLM vision encoders. We identify a consistent three-phase structure of visual processing—diffusion, focus, and rediffusion—and show that the focus phase is where attention most clearly separates strongly and weakly supported visual tokens. However, low attention does not necessarily imply negligible downstream influence: low-attention tokens in the focus phase can still exert non-negligible value-side influence on the attention output relative to their small attention mass. Through controlled phase-wise value interventions, we find that hallucination behavior is particularly sensitive to the value content of low-attention tokens during this phase. Replacing or neutralizing these values reduces hallucination metrics while largely preserving grounded object evidence. A token-level teacher-forcing analysis further shows that the intervention reduces the probability of hallucinated object tokens with much smaller effects on ground-truth object tokens. In addition, Visual Attention Ratio (VAR) analysis shows that focus-phase intervention is accompanied by increased attention to visual tokens during decoding. Based on these observations, we instantiate a simple training-free inference-time intervention that replaces focus-phase low-attention values with an image-level mean value vector using statistics from a single forward pass. Experiments across multiple LVLM backbones demonstrate that this analysis-derived intervention reduces object hallucination with negligible additional runtime and remains compatible with existing inference-time mitigation methods.

Analysis


Hover over each analysis to explore our findings

01 · Three-Phase Structure

Three-Phase Attention Structure in Vision Encoders

Three-phase attention structure across vision-encoder layers

We analyze how attention distributions evolve across vision-encoder layers and find a consistent depth-wise pattern across multiple LVLM backbones, independent of architecture or scale.

We summarize layer-wise attention concentration as \(R^{(\ell)} = \mathbb{E}_h[M^{(\ell,h)}] / \mathbb{E}_h[H^{(\ell,h)}]\), the ratio of the maximum attention score to attention entropy. Tracking \(R^{(\ell)}\) across layers reveals three distinct phases:

  • Phase 1 · Diffusion: in early layers, attention is broadly distributed across visual tokens.
  • Phase 2 · Focus: in intermediate layers, attention becomes sharply concentrated on a small subset of tokens, most clearly separating strongly and weakly supported tokens.
  • Phase 3 · Rediffusion: in later layers, the concentrated pattern spreads out again.

Quantitative Results


Experimental Setup

Value Interventions


We evaluate four LVLM backbones (LLaVA-1.5-7B, LLaVA-1.5-13B, Qwen-2.5-VL, InternVL-2.5) on CHAIR, POPE and MME, using greedy decoding. Following the analysis in the previous section, our intervention replaces the value vectors of the bottom-25% low-attention tokens during the focus phase with an image-level mean value vector, computed from a single forward pass. No training, gradient updates, or iterative optimization is required.

Quantitative

Hallucination Mitigation & Inference Efficiency


We compare against the original models and AUE, which estimates unreliable visual tokens via iterative PGD-based optimization.

  • Consistent hallucination reduction: lowers \(\text{CHAIR}_S\) and \(\text{CHAIR}_I\) across all four backbones, comparable to AUE while generally preserving caption fidelity (F1) better.
  • Preserved visual recognition: maintains or improves F1 and POPE accuracy across Random, Popular, and Adversarial splits, indicating that grounded visual evidence is retained.
  • Negligible runtime overhead: on Qwen-2.5-VL, AUE inflates inference from 2.60s to 33.46s (~12.9× slower); ours runs in 2.71s. On InternVL-2.5, AUE takes 41.29s vs. our 2.44s (~16.9× slower).
CHAIR and POPE benchmark results
Quantitative

General Capability on MME


Beyond hallucination-specific benchmarks, we check on MME whether the intervention disturbs broader multimodal capabilities. We compare per-category MME scores on Qwen-2.5-VL and InternVL-2.5.

  • Profile largely preserved: the radar shapes follow the original models across both perception-oriented categories (existence, count, position, color, OCR, scene, etc.) and cognition-oriented categories (commonsense reasoning, numerical calculation, text translation, code reasoning).
  • No global suppression of visual signals: hallucination reduction on CHAIR and POPE is not achieved at the cost of degraded general multimodal performance.
MME benchmark results

Qualitative Results


Pick a benchmark and a backbone to compare
Benchmark
Model
Qualitative result

BibTeX