This repository demonstrates a research prototype for detecting AI-generated forgeries in videos and images.
The project is built entirely with public datasets and open-source tools, and is designed for educational and research purposes.
It does not contain any proprietary code, data, or deployment logic from any employer.
The rise of AI-generated deepfakes poses new challenges for media authenticity.
This project explores a computer vision pipeline for detecting visual forgeries, focusing on image-based detection from video content.
- Model: CNN-based backbone (ResNet50 variant) fine-tuned on public datasets
- Performance: Achieved ~90–95% F1-score on a benchmark dataset
- Techniques:
- t-SNE embedding visualization for feature analysis
- Threshold tuning based on score distributions
- ROC / PR curve evaluation
- Tools: PyTorch, HuggingFace, OpenCV, Weights & Biases (for experiment tracking)
- Application Scope: Research and academic exploration of video forgery detection
-
Data Preprocessing
- Extract frames from
.mp4videos using parallelized frame extraction scripts - Detect and crop faces using open-source face detection models (e.g., MTCNN, MediaPipe)
- Extract frames from
-
Model Training
- Fine-tune a CNN-based backbone using HuggingFace Trainer API
- Use W&B to log metrics, compare experiments, and tune hyperparameters
-
Evaluation
- Visualize learned embeddings with t-SNE (based on public/synthetic data) to identify clustering patterns
- Adjust classification thresholds based on ROC and PR curve analysis
| File / Folder | Description(Only put model training sample script) |
|---|---|
model_training.ipynb |
CNN fine-tuning with HuggingFace Trainer API |
requirements.txt |
Required Python packages |
This repository is an independent reimplementation for research demonstration.
All data is public or synthetic.
It does not include proprietary datasets, deployment scripts, or configurations from any employer.
Any similarity to real-world systems is coincidental.
- Evaluate on multiple public deepfake datasets (FaceForensics++, Celeb-DF, DFDC)
- Explore ensemble architectures for improved robustness
- Investigate temporal features for frame-sequence-based detection