Three architectures, three data topologies, one question: how does a network's structure have to match the structure of its data? Tabular data with an MLP, images with LeNet-5, sequences with a GRU encoder-decoder — implemented in PyTorch, with convolution and pooling also written by hand and verified against the library.
🇫🇷 Lire ce document en français
End-of-module project, academic year 2025–2026. Full write-up: report/Rapport_Scientifique.pdf (Markdown version).
Dataset: Breast Cancer Wisconsin (569 samples, 30 numeric features). Strict 70 / 15 / 15 train-validation-test split, StandardScaler fitted on the training set only to avoid leakage.
Architecture: 30 → 16 → 8 → 1, ReLU activations, BCEWithLogitsLoss. Built twice — once with nn.Sequential, once as a custom nn.Module — and shown to be equivalent.
| Strategy | Final validation loss | Outcome |
|---|---|---|
| Gaussian N(0, 0.1²) | 0.0243 | Fast convergence |
| Xavier uniform | 0.1213 | Converges, preserves activation variance |
| Constant (0.5) | 0.7503 | Fails |
The constant initialization fails because of neuron symmetry: every neuron computes the same gradient and they all evolve identically, so the layer never differentiates. This is the clearest experimental result of Part I.
| Metric | Value |
|---|---|
| Accuracy | 96.51% |
| Precision | 98.11% |
| Recall | 96.30% |
| F1-score | 97.20% |
Limits. The MLP has a very weak inductive bias — it assumes no spatial or temporal relationship between features. On small tabular datasets it overfits easily and is routinely beaten by Random Forest or XGBoost.
Dataset: MNIST, 28×28 grayscale.
2D cross-correlation and 2×2 max pooling were implemented manually, then compared to nn.Conv2d and nn.MaxPool2d. Maximum difference: 0. The manual implementation is exact, which validates the dimensional reasoning before relying on the library.
| Model | Test accuracy |
|---|---|
| LeNet-5 (CNN) | 97.19% |
| Simple MLP | 96.83% |
Manual conv vs nn.Conv2d |
max diff = 0 |
The gap is narrow — MNIST is an easy problem with centered digits, so an MLP does fine. The real difference shows in the feature maps: the CNN's first layers clearly extract vertical and horizontal edges, a hierarchy the MLP cannot build because flattening destroys the 2D structure.
Task: English → French translation on a short-sentence corpus (Tatoeba style).
Architecture: GRU encoder-decoder. Trained with teacher forcing at a 50% ratio and gradient clipping at max_norm=1.0 to prevent the gradient explosion that plagues RNNs during backpropagation through time. Training loss reaches 0.0003 at epoch 500.
Both decode the test corpus perfectly, which says more about the corpus than about the methods:
i am cold -> j ai froid <EOS>
they play soccer -> ils jouent au foot <EOS>
On a harder corpus, beam search would explore a set of hypotheses and avoid the local optima that greedy decoding walks into. Here the dataset is too simple to separate them — that is an honest limitation of the experiment, not a result in favour of greedy.
Deep learning adapts to data by choosing what the network is allowed to assume:
| Data | Architecture | Assumption baked into the topology |
|---|---|---|
| Tabular | MLP | Everything connects to everything — no order, no topology |
| Image | CNN | Spatial locality and stationarity — a pattern matters regardless of position |
| Sequence | RNN / Seq2Seq | Temporal causality — the past conditions the future, weights shared across time |
Supervised learning succeeds not because of data volume alone, but because of the structural fit between the network topology and the nature of the signal.
.
├── src/
│ ├── Partie1_MLP.py MLP, initialization comparison, test metrics
│ ├── Partie2_CNN.py manual conv/pooling + LeNet-5 on MNIST
│ ├── Partie3_Seq2Seq.py GRU encoder-decoder, greedy and beam search
│ └── generate_missing_plots.py regenerates the figures
├── outputs/part1|part2|part3/ all figures shown above
├── models/best_mlp.pth best Part I checkpoint
└── report/
├── Rapport_Scientifique.pdf full scientific report
├── Rapport_Scientifique.md same content in Markdown
└── Rapport_Scientifique.tex LaTeX source
pip install -r requirements.txt
python src/Partie1_MLP.py # Breast Cancer Wisconsin, ships with scikit-learn
python src/Partie2_CNN.py # MNIST, downloaded automatically into ./data
python src/Partie3_Seq2Seq.py # built-in mini corpusMNIST is not committed — torchvision.datasets.MNIST downloads it into ./data on first run.





