This repository contains the core implementation accompanying the paper Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model.
The method reformulates multimodal aspect-based sentiment analysis (MABSA) as a generative task: given text, an image, and a target aspect, a multimodal large language model jointly generates the aspect sentiment and a natural-language explanation. An aspect-centered dependency subgraph is textualized and added to the prompt to focus the model on sentiment cues related to the target aspect.
For each input text, the code:
- parses every sentence with spaCy and connects consecutive sentence roots;
- locates the target aspect in the resulting dependency graph;
- retains its ancestors and descendants within a specified pruning depth;
- textualizes retained edges as
head -[relation]-> dependent; and - prompts an MLLM to produce
Sentiment: ... Explanation: ....
The same entry points cover all settings reported in the paper:
| Setting | Arguments |
|---|---|
| Vanilla MLLM | --representation none |
| Full dependency graph | --representation edges --hops full |
| Aspect-centered pruning | --representation edges --hops 1, 2, or 3 |
| Without dependency labels | --representation edges-no-rel |
| CoNLL-U-style textualization | --representation conllu |
.
|-- train.py # LoRA fine-tuning and test evaluation
|-- infer.py # zero-shot or adapter-based inference
|-- generate_explanations.py # explanation-augmented TSV construction
|-- src/explainable_mabsa/
| |-- data.py # dataset schema and image loading
| |-- dependency.py # parsing, pruning, and textualization
| |-- output.py # output parsing and metrics
| |-- prompts.py # prompts from the paper
| |-- runtime.py # model/processor integration
| `-- workflow.py # shared preparation and evaluation
`-- tests/test_core.py
Datasets, images, model weights, logs, predictions, and reported experimental results are intentionally excluded.
Python 3.10 or 3.11 and a CUDA-enabled PyTorch environment are recommended.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m spacy download en_core_web_smThe unified implementation uses the current Hugging Face multimodal model interface. The paper evaluates Qwen3-VL-4B, Qwen3-VL-8B, and Ministral-3-8B; corresponding public checkpoints include:
Qwen/Qwen3-VL-4B-InstructQwen/Qwen3-VL-8B-Instructmistralai/Ministral-3-8B-Instruct-2512
A local checkpoint directory can be passed anywhere a model ID is accepted.
Obtain Twitter2015 or Twitter2017 and their associated images from the original dataset provider. Data is not redistributed here. Arrange each dataset as follows:
/path/to/twitter2015/
|-- train_with_explanations.tsv
|-- val_with_explanations.tsv
`-- test_with_explanations.tsv
/path/to/twitter2015_images/
|-- 1000.jpg
`-- ...
The loader recognizes both the original column names and concise aliases:
| Value | Original header | Alias |
|---|---|---|
| sentiment ID | #1 Label |
label or sentiment |
| image filename | #2 ImageID |
image, image_id, or image_path |
text with optional $T$ placeholder |
#3 String |
text or sentence |
| aspect | #3 String.1 |
aspect, aspect_term, or target |
| explanation | #4 explanation |
explanation or rationale |
Sentiment IDs must be 0 (negative), 1 (neutral), or 2 (positive). Training requires explanations. Inference can use a TSV without the explanation column; in that case only classification metrics are computed.
The paper uses Qwen3-VL-32B to generate explanations while conditioning on gold sentiment labels. This command always writes a new file and refuses to overwrite the input TSV:
python generate_explanations.py \
--model Qwen/Qwen3-VL-32B-Instruct \
--input-tsv /path/to/twitter2015/train.tsv \
--image-dir /path/to/twitter2015_images \
--output-tsv /path/to/twitter2015/train_with_explanations.tsvRun the command separately for the training, validation, and test splits. As described in the paper, manually inspect a random 10% sample before using generated explanations as supervision.
The following reproduces the principal dependency-guided setting with pruning depth 2:
python train.py \
--model Qwen/Qwen3-VL-4B-Instruct \
--data-dir /path/to/twitter2015 \
--image-dir /path/to/twitter2015_images \
--output-dir outputs/qwen3-vl-4b-twitter2015-n2 \
--representation edges \
--hops 2 \
--device cuda:0 \
--bertscore-model FacebookAI/roberta-largeDefaults match the paper: 10 epochs, batch size 1, AdamW with learning rate 1e-5, and LoRA with rank 4, alpha 16, and dropout 0.1. LoRA is applied to q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj.
To run the vanilla baseline or the full-syntax baseline, change only the representation arguments:
# Vanilla
python train.py ... --representation none
# Full dependency graph (n = infinity)
python train.py ... --representation edges --hops fullThe trained LoRA adapter is saved under <output-dir>/adapter/. Predictions and aggregate metrics are written alongside it. These runtime artifacts are ignored by Git.
Zero-shot inference uses the base model directly:
python infer.py \
--model Qwen/Qwen3-VL-4B-Instruct \
--test-tsv /path/to/twitter2015/test_with_explanations.tsv \
--image-dir /path/to/twitter2015_images \
--representation edges \
--hops 2 \
--output outputs/zeroshot-twitter2015.jsonlFor a fine-tuned model, add the saved adapter:
python infer.py \
--model Qwen/Qwen3-VL-4B-Instruct \
--adapter outputs/qwen3-vl-4b-twitter2015-n2/adapter \
--test-tsv /path/to/twitter2015/test_with_explanations.tsv \
--image-dir /path/to/twitter2015_images \
--representation edges \
--hops 2The evaluator reports accuracy and macro-F1 for sentiment classification, and BLEU, ROUGE-1/2/L, and optional BERTScore-F1 for explanations. Pass --bertscore-model FacebookAI/roberta-large to enable the BERTScore configuration used in the paper.
The unit tests exercise dependency pruning, textualization, structured prediction parsing, and pruning-depth handling without requiring model weights or dataset files:
python -m unittest discover -s tests -vPlease cite the paper if this code is useful in your research. Publication metadata and the final BibTeX entry will be added after publication.
Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model
Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, and Yan Song