Skip to content
VisionXLabPublic

About

Code for: Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

arXiv Hugging Face

🔍 Overview

We find that LVLMs exhibit an intrinsic vision-attending tendency, which can serve as a reliable signal for adaptive visual steering. Beyond visual information, we further show that appropriately incorporating prefilled textual context and generation history at different decoding steps also contributes to hallucination mitigation.

Based on these observations, we propose AIMS, a training-free method that adaptively coordinates visual, prefilled, and generated context during decoding.

🛠️ Environment

We use separate environments for the three LVLMs due to their different dependencies.

LLaVA-1.5

conda create -n aims_llava python=3.10 -y
conda activate aims_llava
pip install -r requirements_llava.txt

Download liuhaotian/llava-v1.5-7b from Hugging Face and set the local model path in model_loader.py.

Qwen2.5-VL

conda create -n aims_qwen25vl python=3.10 -y
conda activate aims_qwen25vl
pip install -r requirements_qwen25vl.txt

Download Qwen/Qwen2.5-VL-3B-Instruct from Hugging Face and set the local model path in the corresponding inference script.

Qwen3.5

conda create -n aims_qwen35 python=3.10 -y
conda activate aims_qwen35
pip install -r requirements_qwen35.txt

Download Qwen/Qwen3.5-9B from Hugging Face and set the local model path in the corresponding inference script.

📦 Data Preparation

The benchmarks used in our experiments are available on Hugging Face.

🚀 Inference

CHAIR

cd AIMS

bash scripts_chair/chair_llava15_aims.sh    # LLaVA-1.5
bash scripts_chair/chair_qwen25vl_aims.sh   # Qwen2.5-VL
bash scripts_chair/chair_qwen35_aims.sh     # Qwen3.5

FaithScore

cd AIMS

bash scripts_faithscore/fs_qwen25vl.sh      # Qwen2.5-VL

AMBER-G

cd AIMS

bash scripts_amber/amber_qwen25vl.sh        # Qwen2.5-VL
bash scripts_amber/amber_qwen35.sh          # Qwen3.5

MME

pip install scikit-learn
cd AIMS

bash scripts_mme/mme_qwen25vl_aims.sh       # Qwen2.5-VL
bash scripts_mme/mme_qwen35_aims.sh         # Qwen3.5

⚙️ Key Arguments

AIMS

Argument Description
--use-qsteer-adaptive Enable AIMS during decoding.
--start-layer First decoder layer to apply AIMS.
--end-layer Last decoder layer to apply AIMS.
--alpha Overall steering strength.
--visual-branch Enable steering from visual context.
--visual-sigma Temperature coefficient for the visual branch.
--prefill-branch Enable steering from prefilled textual context.
--prefill-sigma Temperature coefficient for the prefill branch.
--decode-branch Enable steering from previously generated context.
--decode-window Number of previous decoding tokens used by the generated-context branch.
--decode-sigma Temperature coefficient for the generated-context branch.

Decoding

We support three decoding strategies:

  • Greedy decoding: --beam 1
  • Beam search: --beam 5
  • Nucleus sampling: --beam 1 --sample

Other Arguments

Argument Description
--model-path Path to the LVLM checkpoint.
--data-path Path to benchmark images.
--bench-type Evaluation benchmark, e.g., chair.
--exp-tag Optional tag for identifying the experiment.
--debug-number Number of samples used for debugging.

📊 Evaluation

Except for FaithScore, all evaluation scripts can be run directly in the corresponding AIMS environment without creating a separate evaluation environment. FaithScore requires an additional environment due to its specific dependencies.

CHAIR

The required NLTK 3.8.1 data can be downloaded from Hugging Face:

VisionXLab/AIMS_Benchmarks/nltk_3-8-1

After downloading, set NLTK_DATA to the local path of nltk_3-8-1 in chair.py:

NLTK_DATA = "<path_to_nltk_3-8-1>"

Then run:

bash scripts/chair_score.sh

FaithScore

For environment setup and evaluation scripts, please refer to FAITHSCORE/README.md.

AMBER-G

The required NLTK 3.8.1 data can also be downloaded from Hugging Face:

sharon11/aims_benchmarks/nltk_3-8-1

After downloading, set NLTK_DATA to the local path of nltk_3-8-1 in AMBER/inference.py:

NLTK_DATA = "<path_to_nltk_3-8-1>"

Install the additional dependencies and run the evaluation:

cd AMBER

pip install spacy
python -m spacy download en_core_web_lg

bash scripts/run.sh

Alternatively, en_core_web_lg can be installed manually:

# Download:
# https://github.com/explosion/spacy-models/releases/download/en_core_web_lg-3.8.0/en_core_web_lg-3.8.0-py3-none-any.whl

python -m pip install en_core_web_lg-3.8.0-py3-none-any.whl

MME

pip install scikit-learn
bash scripts/mme_score.sh

🙏 Acknowledgements

We sincerely thank the authors and contributors of the following projects and benchmarks:

📚 Citation

@article{aims,
  title   = {AIMS: Adaptive Information Multi-source Steering for Hallucination Mitigation in Large Vision-Language Models},
  author  = {...},
  journal = {arXiv preprint},
  year    = {2026}
}

About

Code for: Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages