Skip to content

Repository files navigation

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

This repository contains the core implementation accompanying the paper Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model.

The method reformulates multimodal aspect-based sentiment analysis (MABSA) as a generative task: given text, an image, and a target aspect, a multimodal large language model jointly generates the aspect sentiment and a natural-language explanation. An aspect-centered dependency subgraph is textualized and added to the prompt to focus the model on sentiment cues related to the target aspect.

Method

For each input text, the code:

  1. parses every sentence with spaCy and connects consecutive sentence roots;
  2. locates the target aspect in the resulting dependency graph;
  3. retains its ancestors and descendants within a specified pruning depth;
  4. textualizes retained edges as head -[relation]-> dependent; and
  5. prompts an MLLM to produce Sentiment: ... Explanation: ....

The same entry points cover all settings reported in the paper:

Setting Arguments
Vanilla MLLM --representation none
Full dependency graph --representation edges --hops full
Aspect-centered pruning --representation edges --hops 1, 2, or 3
Without dependency labels --representation edges-no-rel
CoNLL-U-style textualization --representation conllu

Repository Layout

.
|-- train.py                  # LoRA fine-tuning and test evaluation
|-- infer.py                  # zero-shot or adapter-based inference
|-- generate_explanations.py  # explanation-augmented TSV construction
|-- src/explainable_mabsa/
|   |-- data.py               # dataset schema and image loading
|   |-- dependency.py         # parsing, pruning, and textualization
|   |-- output.py             # output parsing and metrics
|   |-- prompts.py            # prompts from the paper
|   |-- runtime.py            # model/processor integration
|   `-- workflow.py           # shared preparation and evaluation
`-- tests/test_core.py

Datasets, images, model weights, logs, predictions, and reported experimental results are intentionally excluded.

Installation

Python 3.10 or 3.11 and a CUDA-enabled PyTorch environment are recommended.

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m spacy download en_core_web_sm

The unified implementation uses the current Hugging Face multimodal model interface. The paper evaluates Qwen3-VL-4B, Qwen3-VL-8B, and Ministral-3-8B; corresponding public checkpoints include:

  • Qwen/Qwen3-VL-4B-Instruct
  • Qwen/Qwen3-VL-8B-Instruct
  • mistralai/Ministral-3-8B-Instruct-2512

A local checkpoint directory can be passed anywhere a model ID is accepted.

Data Preparation

Obtain Twitter2015 or Twitter2017 and their associated images from the original dataset provider. Data is not redistributed here. Arrange each dataset as follows:

/path/to/twitter2015/
|-- train_with_explanations.tsv
|-- val_with_explanations.tsv
`-- test_with_explanations.tsv

/path/to/twitter2015_images/
|-- 1000.jpg
`-- ...

The loader recognizes both the original column names and concise aliases:

Value Original header Alias
sentiment ID #1 Label label or sentiment
image filename #2 ImageID image, image_id, or image_path
text with optional $T$ placeholder #3 String text or sentence
aspect #3 String.1 aspect, aspect_term, or target
explanation #4 explanation explanation or rationale

Sentiment IDs must be 0 (negative), 1 (neutral), or 2 (positive). Training requires explanations. Inference can use a TSV without the explanation column; in that case only classification metrics are computed.

Construct Explanation-Augmented Data

The paper uses Qwen3-VL-32B to generate explanations while conditioning on gold sentiment labels. This command always writes a new file and refuses to overwrite the input TSV:

python generate_explanations.py \
  --model Qwen/Qwen3-VL-32B-Instruct \
  --input-tsv /path/to/twitter2015/train.tsv \
  --image-dir /path/to/twitter2015_images \
  --output-tsv /path/to/twitter2015/train_with_explanations.tsv

Run the command separately for the training, validation, and test splits. As described in the paper, manually inspect a random 10% sample before using generated explanations as supervision.

Fine-Tuning

The following reproduces the principal dependency-guided setting with pruning depth 2:

python train.py \
  --model Qwen/Qwen3-VL-4B-Instruct \
  --data-dir /path/to/twitter2015 \
  --image-dir /path/to/twitter2015_images \
  --output-dir outputs/qwen3-vl-4b-twitter2015-n2 \
  --representation edges \
  --hops 2 \
  --device cuda:0 \
  --bertscore-model FacebookAI/roberta-large

Defaults match the paper: 10 epochs, batch size 1, AdamW with learning rate 1e-5, and LoRA with rank 4, alpha 16, and dropout 0.1. LoRA is applied to q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj.

To run the vanilla baseline or the full-syntax baseline, change only the representation arguments:

# Vanilla
python train.py ... --representation none

# Full dependency graph (n = infinity)
python train.py ... --representation edges --hops full

The trained LoRA adapter is saved under <output-dir>/adapter/. Predictions and aggregate metrics are written alongside it. These runtime artifacts are ignored by Git.

Inference

Zero-shot inference uses the base model directly:

python infer.py \
  --model Qwen/Qwen3-VL-4B-Instruct \
  --test-tsv /path/to/twitter2015/test_with_explanations.tsv \
  --image-dir /path/to/twitter2015_images \
  --representation edges \
  --hops 2 \
  --output outputs/zeroshot-twitter2015.jsonl

For a fine-tuned model, add the saved adapter:

python infer.py \
  --model Qwen/Qwen3-VL-4B-Instruct \
  --adapter outputs/qwen3-vl-4b-twitter2015-n2/adapter \
  --test-tsv /path/to/twitter2015/test_with_explanations.tsv \
  --image-dir /path/to/twitter2015_images \
  --representation edges \
  --hops 2

The evaluator reports accuracy and macro-F1 for sentiment classification, and BLEU, ROUGE-1/2/L, and optional BERTScore-F1 for explanations. Pass --bertscore-model FacebookAI/roberta-large to enable the BERTScore configuration used in the paper.

Tests

The unit tests exercise dependency pruning, textualization, structured prediction parsing, and pruning-depth handling without requiring model weights or dataset files:

python -m unittest discover -s tests -v

Citation

Please cite the paper if this code is useful in your research. Publication metadata and the final BibTeX entry will be added after publication.

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model
Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, and Yan Song

About

This repository contains the core implementation accompanying the paper Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages