A small, fully offline research project testing a generalization failure mode in adapter-based self-interpretation methods (à la AI Alignment Foundation's Self-Interpretation via Adapter Probes).
Full write-up: see WRITEUP.md.
A full-rank adapter can hit 93% self-report accuracy on concepts it was trained on, but only ~11% on concepts it never saw a label for — even though the transformation it needs to invert is global and concept-agnostic, so a correct solution should generalize perfectly. A mechanistic check shows the adapter only partially recovers the true underlying transformation (cosine similarity 0.30 with the ground-truth inverse), consistent with shortcut learning rather than genuine recovery of model structure.
Real interpretability evaluation is expensive and slow to compare against (you rarely have a "ground truth" transformation to check an adapter against). This project builds a synthetic toy world where the ground truth is known by construction, so adapter behavior can be checked mechanistically, not just by accuracy — the same way toy models of superposition let you check "does this SAE really recover the true features" when you built the features yourself.
- 12 synthetic "concept directions" in a 64-dim space
- Model activations = concept signal passed through a fixed, global, invertible linear distortion (rotation + shear) + noise
- Monosemantic (1 concept active) and polysemantic (2-3 concepts active) regimes
- 4 of 12 concepts held out entirely from adapter training
- Three adapter parameterizations: bias-only, scalar-affine, full-rank (matching the sweep in the real adapter-probes paper)
- Fixed cosine-similarity readout head (adapter's only job: reshape the activation so this fixed readout works)
pip install torch numpy scikit-learn matplotlib
python run_experiment.py # main accuracy table -> results.json
python diagnostics.py # mechanistic check + generalization sweep -> diagnostics_results.jsonBoth scripts run in well under a minute on CPU. No downloaded weights, no API keys, no internet access required — fully self-contained and reproducible via fixed seeds.
concepts.py # synthetic activation-space generator (ConceptWorld)
adapters.py # bias-only / scalar-affine / full-rank adapter modules
run_experiment.py # main experiment: train + eval all adapters across 4 conditions
diagnostics.py # mechanistic check (W vs true M^-1) + training-breadth sweep
results_plot.png # summary figure
WRITEUP.md # full LessWrong-style report of motivation/method/results/limitations
See WRITEUP.md — main ones: fully linear/synthetic toy setting, single random seed for the main table, fixed (not trainable) readout head, and the polysemantic condition isn't fully isolated from the seen/held-out split.
