We find that LVLMs exhibit an intrinsic vision-attending tendency, which can serve as a reliable signal for adaptive visual steering. Beyond visual information, we further show that appropriately incorporating prefilled textual context and generation history at different decoding steps also contributes to hallucination mitigation.
Based on these observations, we propose AIMS, a training-free method that adaptively coordinates visual, prefilled, and generated context during decoding.
We use separate environments for the three LVLMs due to their different dependencies.
conda create -n aims_llava python=3.10 -y
conda activate aims_llava
pip install -r requirements_llava.txtDownload liuhaotian/llava-v1.5-7b from Hugging Face and set the local model path in model_loader.py.
conda create -n aims_qwen25vl python=3.10 -y
conda activate aims_qwen25vl
pip install -r requirements_qwen25vl.txtDownload Qwen/Qwen2.5-VL-3B-Instruct from Hugging Face and set the local model path in the corresponding inference script.
conda create -n aims_qwen35 python=3.10 -y
conda activate aims_qwen35
pip install -r requirements_qwen35.txtDownload Qwen/Qwen3.5-9B from Hugging Face and set the local model path in the corresponding inference script.
The benchmarks used in our experiments are available on Hugging Face.
cd AIMS
bash scripts_chair/chair_llava15_aims.sh # LLaVA-1.5
bash scripts_chair/chair_qwen25vl_aims.sh # Qwen2.5-VL
bash scripts_chair/chair_qwen35_aims.sh # Qwen3.5cd AIMS
bash scripts_faithscore/fs_qwen25vl.sh # Qwen2.5-VLcd AIMS
bash scripts_amber/amber_qwen25vl.sh # Qwen2.5-VL
bash scripts_amber/amber_qwen35.sh # Qwen3.5pip install scikit-learn
cd AIMS
bash scripts_mme/mme_qwen25vl_aims.sh # Qwen2.5-VL
bash scripts_mme/mme_qwen35_aims.sh # Qwen3.5| Argument | Description |
|---|---|
--use-qsteer-adaptive |
Enable AIMS during decoding. |
--start-layer |
First decoder layer to apply AIMS. |
--end-layer |
Last decoder layer to apply AIMS. |
--alpha |
Overall steering strength. |
--visual-branch |
Enable steering from visual context. |
--visual-sigma |
Temperature coefficient for the visual branch. |
--prefill-branch |
Enable steering from prefilled textual context. |
--prefill-sigma |
Temperature coefficient for the prefill branch. |
--decode-branch |
Enable steering from previously generated context. |
--decode-window |
Number of previous decoding tokens used by the generated-context branch. |
--decode-sigma |
Temperature coefficient for the generated-context branch. |
We support three decoding strategies:
- Greedy decoding:
--beam 1 - Beam search:
--beam 5 - Nucleus sampling:
--beam 1 --sample
| Argument | Description |
|---|---|
--model-path |
Path to the LVLM checkpoint. |
--data-path |
Path to benchmark images. |
--bench-type |
Evaluation benchmark, e.g., chair. |
--exp-tag |
Optional tag for identifying the experiment. |
--debug-number |
Number of samples used for debugging. |
Except for FaithScore, all evaluation scripts can be run directly in the corresponding AIMS environment without creating a separate evaluation environment. FaithScore requires an additional environment due to its specific dependencies.
The required NLTK 3.8.1 data can be downloaded from Hugging Face:
VisionXLab/AIMS_Benchmarks/nltk_3-8-1
After downloading, set NLTK_DATA to the local path of nltk_3-8-1 in chair.py:
NLTK_DATA = "<path_to_nltk_3-8-1>"Then run:
bash scripts/chair_score.shFor environment setup and evaluation scripts, please refer to FAITHSCORE/README.md.
The required NLTK 3.8.1 data can also be downloaded from Hugging Face:
sharon11/aims_benchmarks/nltk_3-8-1
After downloading, set NLTK_DATA to the local path of nltk_3-8-1 in AMBER/inference.py:
NLTK_DATA = "<path_to_nltk_3-8-1>"Install the additional dependencies and run the evaluation:
cd AMBER
pip install spacy
python -m spacy download en_core_web_lg
bash scripts/run.shAlternatively, en_core_web_lg can be installed manually:
# Download:
# https://github.com/explosion/spacy-models/releases/download/en_core_web_lg-3.8.0/en_core_web_lg-3.8.0-py3-none-any.whl
python -m pip install en_core_web_lg-3.8.0-py3-none-any.whlpip install scikit-learn
bash scripts/mme_score.shWe sincerely thank the authors and contributors of the following projects and benchmarks:
@article{aims,
title = {AIMS: Adaptive Information Multi-source Steering for Hallucination Mitigation in Large Vision-Language Models},
author = {...},
journal = {arXiv preprint},
year = {2026}
}