Minor Project – B.Tech (CSE)
Indian Institute of Information Technology, Bhopal
April 2026
Memes are a powerful form of online communication that combine visual content and overlaid text to express humor, sarcasm, emotion, and sometimes harmful intent. Detecting cyberbullying or offensive content in memes is more challenging than analyzing plain text because meaning emerges from the interaction between image and text.
This project presents a Multimodal Multi-Task Learning framework for meme analysis, focusing on four interconnected tasks:
- Cyberbullying Detection
- Sentiment Classification
- Emotion Recognition
- Sarcasm Detection
We compare image-only, text-only, and multimodal architectures to evaluate how combining modalities improves meme understanding.
- Comparative study of unimodal vs multimodal approaches
- Unified multi-task learning framework for four related classification tasks
- ResNet-18 based visual encoder
- Transformer-based text encoder for OCR-extracted meme text
- Intermediate fusion-based multimodal architecture
- Detailed evaluation using Accuracy, Precision, Recall, and Macro F1-score
- Experimental analysis with training curves and task-wise comparisons
- Backbone: ResNet-18
- 512-dimensional feature vector
- Four task-specific classification heads
- OCR-extracted meme text
- Transformer-based encoder (BERT-style)
- Shared embedding with four task heads
- Text Encoder + Image Encoder
- Feature Concatenation (Intermediate Fusion)
- Fusion MLP
- Four task-specific heads
Pipeline:
Meme → (Image Encoder + Text Encoder)
→ Feature Fusion
→ Shared Representation
→ Bully | Sentiment | Emotion | Sarcasm
After preprocessing:
- Training: 3,474 memes
- Validation: 738 memes
- Testing: 726 memes
Each meme contains:
- Image file
- OCR-extracted text
- Four labels (bully, sentiment, emotion, sarcasm)
Preprocessing steps:
- Text cleaning
- Label validation
- Image verification
- Label encoding
- PyTorch Dataset and DataLoader creation
Python 3.8+
PyTorch
Torchvision
HuggingFace Transformers
Scikit-learn
Pandas / NumPy
Matplotlib / Seaborn
CUDA (for GPU acceleration)
- Optimizer: AdamW
- Batch Size: 16
- Image Size: 224 × 224
- Dropout: 0.55
- Weight Decay: 3e-4
- Early Stopping Patience: 4 epochs
- Loss: Cross-Entropy / Focal Loss with class balancing
Total Loss:
L = αL_bully + βL_sentiment + γL_emotion + δL_sarcasm
Overall Test Performance:
Image-only model
Average Accuracy: 66.46%
Average Macro-F1: 65.18%
Text-only model
Average Accuracy: 68.46%
Average Macro-F1: 69.24%
Multimodal model
Average Accuracy: 71.32%
Average Macro-F1: 70.79%
Best Task Performance (Multimodal):
Bullying Detection: 87.05%
Sentiment Classification: 74.24%
Emotion Recognition: 51.65%
Sarcasm Detection: 72.31%
The multimodal model significantly improves bully and sarcasm detection, demonstrating complementary information between image and text modalities.
- Image-only model performs well for bullying detection
- Text-only model excels in sentiment analysis
- Emotion detection is the hardest task (10 classes + imbalance)
- Multimodal learning provides the most robust overall performance
- Regularization techniques reduce overfitting
- Multimodal Transformers (CLIP, ViLT, VisualBERT)
- Cross-modal attention mechanisms
- Advanced imbalance handling (SMOTE, GANs)
- Grad-CAM and SHAP explainability
- Extension to other code-mixed languages
- Deployment as API or browser moderation tool
Vasudev Gupta (23U02081)
Harshvardhan Singh (23U02091)
Under the supervision of
Dr. Yatendra Sahu
Assistant Professor, Department of CSE
IIIT Bhopal
Inspired by:
Maity et al., 2022 – EMPIRICAL: Emotion Aware Multimodal Multitask Learning for Cyberbullying Detection in Code-Mixed Memes (COLING 2022)
Developed for academic and research purposes.
Free to use for educational or research work with proper citation.
If you found this project useful, consider giving it a star ⭐