Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

🧠 Study of Multimodal Multi-Task Learning for Meme Analysis

Minor Project – B.Tech (CSE)
Indian Institute of Information Technology, Bhopal
April 2026


📌 Overview

Memes are a powerful form of online communication that combine visual content and overlaid text to express humor, sarcasm, emotion, and sometimes harmful intent. Detecting cyberbullying or offensive content in memes is more challenging than analyzing plain text because meaning emerges from the interaction between image and text.

This project presents a Multimodal Multi-Task Learning framework for meme analysis, focusing on four interconnected tasks:

  • Cyberbullying Detection
  • Sentiment Classification
  • Emotion Recognition
  • Sarcasm Detection

We compare image-only, text-only, and multimodal architectures to evaluate how combining modalities improves meme understanding.


🎯 Key Contributions

  • Comparative study of unimodal vs multimodal approaches
  • Unified multi-task learning framework for four related classification tasks
  • ResNet-18 based visual encoder
  • Transformer-based text encoder for OCR-extracted meme text
  • Intermediate fusion-based multimodal architecture
  • Detailed evaluation using Accuracy, Precision, Recall, and Macro F1-score
  • Experimental analysis with training curves and task-wise comparisons

🏗️ System Architecture

1) Image-Only Model

  • Backbone: ResNet-18
  • 512-dimensional feature vector
  • Four task-specific classification heads

2) Text-Only Model

  • OCR-extracted meme text
  • Transformer-based encoder (BERT-style)
  • Shared embedding with four task heads

3) Multimodal Model (Final Model)

  • Text Encoder + Image Encoder
  • Feature Concatenation (Intermediate Fusion)
  • Fusion MLP
  • Four task-specific heads

Pipeline:

Meme → (Image Encoder + Text Encoder)
→ Feature Fusion
→ Shared Representation
→ Bully | Sentiment | Emotion | Sarcasm


📂 Dataset Details

After preprocessing:

  • Training: 3,474 memes
  • Validation: 738 memes
  • Testing: 726 memes

Each meme contains:

  • Image file
  • OCR-extracted text
  • Four labels (bully, sentiment, emotion, sarcasm)

Preprocessing steps:

  • Text cleaning
  • Label validation
  • Image verification
  • Label encoding
  • PyTorch Dataset and DataLoader creation

⚙️ Tech Stack

Python 3.8+
PyTorch
Torchvision
HuggingFace Transformers
Scikit-learn
Pandas / NumPy
Matplotlib / Seaborn
CUDA (for GPU acceleration)


🔬 Training Configuration

  • Optimizer: AdamW
  • Batch Size: 16
  • Image Size: 224 × 224
  • Dropout: 0.55
  • Weight Decay: 3e-4
  • Early Stopping Patience: 4 epochs
  • Loss: Cross-Entropy / Focal Loss with class balancing

Total Loss:

L = αL_bully + βL_sentiment + γL_emotion + δL_sarcasm


📊 Experimental Results

Overall Test Performance:

Image-only model
Average Accuracy: 66.46%
Average Macro-F1: 65.18%

Text-only model
Average Accuracy: 68.46%
Average Macro-F1: 69.24%

Multimodal model
Average Accuracy: 71.32%
Average Macro-F1: 70.79%

Best Task Performance (Multimodal):

Bullying Detection: 87.05%
Sentiment Classification: 74.24%
Emotion Recognition: 51.65%
Sarcasm Detection: 72.31%

The multimodal model significantly improves bully and sarcasm detection, demonstrating complementary information between image and text modalities.


📉 Key Observations

  • Image-only model performs well for bullying detection
  • Text-only model excels in sentiment analysis
  • Emotion detection is the hardest task (10 classes + imbalance)
  • Multimodal learning provides the most robust overall performance
  • Regularization techniques reduce overfitting

🚀 Future Improvements

  • Multimodal Transformers (CLIP, ViLT, VisualBERT)
  • Cross-modal attention mechanisms
  • Advanced imbalance handling (SMOTE, GANs)
  • Grad-CAM and SHAP explainability
  • Extension to other code-mixed languages
  • Deployment as API or browser moderation tool

👨‍💻 Authors

Vasudev Gupta (23U02081)
Harshvardhan Singh (23U02091)

Under the supervision of
Dr. Yatendra Sahu
Assistant Professor, Department of CSE
IIIT Bhopal


📚 Reference

Inspired by:
Maity et al., 2022 – EMPIRICAL: Emotion Aware Multimodal Multitask Learning for Cyberbullying Detection in Code-Mixed Memes (COLING 2022)


📝 License

Developed for academic and research purposes.
Free to use for educational or research work with proper citation.


If you found this project useful, consider giving it a star ⭐

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages