Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Deep Learning — MLP, CNN and Seq2Seq from the ground up

Three architectures, three data topologies, one question: how does a network's structure have to match the structure of its data? Tabular data with an MLP, images with LeNet-5, sequences with a GRU encoder-decoder — implemented in PyTorch, with convolution and pooling also written by hand and verified against the library.

Python PyTorch scikit-learn License: MIT

🇫🇷 Lire ce document en français

End-of-module project, academic year 2025–2026. Full write-up: report/Rapport_Scientifique.pdf (Markdown version).


Part I — Multilayer Perceptron on tabular data

Dataset: Breast Cancer Wisconsin (569 samples, 30 numeric features). Strict 70 / 15 / 15 train-validation-test split, StandardScaler fitted on the training set only to avoid leakage.

Architecture: 30 → 16 → 8 → 1, ReLU activations, BCEWithLogitsLoss. Built twice — once with nn.Sequential, once as a custom nn.Module — and shown to be equivalent.

Weight initialization decides whether the network learns at all

Strategy Final validation loss Outcome
Gaussian N(0, 0.1²) 0.0243 Fast convergence
Xavier uniform 0.1213 Converges, preserves activation variance
Constant (0.5) 0.7503 Fails

The constant initialization fails because of neuron symmetry: every neuron computes the same gradient and they all evolve identically, so the layer never differentiates. This is the clearest experimental result of Part I.

Test set performance

Metric Value
Accuracy 96.51%
Precision 98.11%
Recall 96.30%
F1-score 97.20%

Validation loss curves for the three weight initialization strategies Confusion matrix on the test set

Limits. The MLP has a very weak inductive bias — it assumes no spatial or temporal relationship between features. On small tabular datasets it overfits easily and is routinely beaten by Random Forest or XGBoost.


Part II — Convolutional networks on images

Dataset: MNIST, 28×28 grayscale.

Hand-written convolution, verified against PyTorch

2D cross-correlation and 2×2 max pooling were implemented manually, then compared to nn.Conv2d and nn.MaxPool2d. Maximum difference: 0. The manual implementation is exact, which validates the dimensional reasoning before relying on the library.

LeNet-5 vs a plain MLP

Model Test accuracy
LeNet-5 (CNN) 97.19%
Simple MLP 96.83%
Manual conv vs nn.Conv2d max diff = 0

The gap is narrow — MNIST is an easy problem with centered digits, so an MLP does fine. The real difference shows in the feature maps: the CNN's first layers clearly extract vertical and horizontal edges, a hierarchy the MLP cannot build because flattening destroys the 2D structure.

Training loss, LeNet-5 versus MLP on MNIST First-layer feature maps showing edge detectors


Part III — Sequence models: GRU encoder-decoder

Task: English → French translation on a short-sentence corpus (Tatoeba style).

Architecture: GRU encoder-decoder. Trained with teacher forcing at a 50% ratio and gradient clipping at max_norm=1.0 to prevent the gradient explosion that plagues RNNs during backpropagation through time. Training loss reaches 0.0003 at epoch 500.

Greedy decoding vs beam search

Both decode the test corpus perfectly, which says more about the corpus than about the methods:

i am cold        ->  j ai froid <EOS>
they play soccer ->  ils jouent au foot <EOS>

On a harder corpus, beam search would explore a set of hypotheses and avoid the local optima that greedy decoding walks into. Here the dataset is too simple to separate them — that is an honest limitation of the experiment, not a result in favour of greedy.

Seq2Seq training loss over 500 epochs Gradient norm across training, showing the effect of clipping


The transversal question: inductive bias

Deep learning adapts to data by choosing what the network is allowed to assume:

Data Architecture Assumption baked into the topology
Tabular MLP Everything connects to everything — no order, no topology
Image CNN Spatial locality and stationarity — a pattern matters regardless of position
Sequence RNN / Seq2Seq Temporal causality — the past conditions the future, weights shared across time

Supervised learning succeeds not because of data volume alone, but because of the structural fit between the network topology and the nature of the signal.


Repository layout

.
├── src/
│   ├── Partie1_MLP.py            MLP, initialization comparison, test metrics
│   ├── Partie2_CNN.py            manual conv/pooling + LeNet-5 on MNIST
│   ├── Partie3_Seq2Seq.py        GRU encoder-decoder, greedy and beam search
│   └── generate_missing_plots.py regenerates the figures
├── outputs/part1|part2|part3/    all figures shown above
├── models/best_mlp.pth           best Part I checkpoint
└── report/
    ├── Rapport_Scientifique.pdf  full scientific report
    ├── Rapport_Scientifique.md   same content in Markdown
    └── Rapport_Scientifique.tex  LaTeX source

Running it

pip install -r requirements.txt

python src/Partie1_MLP.py       # Breast Cancer Wisconsin, ships with scikit-learn
python src/Partie2_CNN.py       # MNIST, downloaded automatically into ./data
python src/Partie3_Seq2Seq.py   # built-in mini corpus

MNIST is not committed — torchvision.datasets.MNIST downloads it into ./data on first run.

License

MIT

About

MLP, LeNet-5 CNN and GRU Seq2Seq built in PyTorch, with convolution and pooling implemented by hand and verified against the library. Includes the full scientific report and all result figures.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages