Skip to content

About

A small, fully offline research project testing a generalization failure mode in adapter-based self-interpretation methods

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Self-Interpretation Adapters: A Shortcut-Learning Stress Test

A small, fully offline research project testing a generalization failure mode in adapter-based self-interpretation methods (à la AI Alignment Foundation's Self-Interpretation via Adapter Probes).

Full write-up: see WRITEUP.md.

TL;DR

A full-rank adapter can hit 93% self-report accuracy on concepts it was trained on, but only ~11% on concepts it never saw a label for — even though the transformation it needs to invert is global and concept-agnostic, so a correct solution should generalize perfectly. A mechanistic check shows the adapter only partially recovers the true underlying transformation (cosine similarity 0.30 with the ground-truth inverse), consistent with shortcut learning rather than genuine recovery of model structure.

Why this setup

Real interpretability evaluation is expensive and slow to compare against (you rarely have a "ground truth" transformation to check an adapter against). This project builds a synthetic toy world where the ground truth is known by construction, so adapter behavior can be checked mechanistically, not just by accuracy — the same way toy models of superposition let you check "does this SAE really recover the true features" when you built the features yourself.

Setup

  • 12 synthetic "concept directions" in a 64-dim space
  • Model activations = concept signal passed through a fixed, global, invertible linear distortion (rotation + shear) + noise
  • Monosemantic (1 concept active) and polysemantic (2-3 concepts active) regimes
  • 4 of 12 concepts held out entirely from adapter training
  • Three adapter parameterizations: bias-only, scalar-affine, full-rank (matching the sweep in the real adapter-probes paper)
  • Fixed cosine-similarity readout head (adapter's only job: reshape the activation so this fixed readout works)

Run it

pip install torch numpy scikit-learn matplotlib
python run_experiment.py      # main accuracy table -> results.json
python diagnostics.py         # mechanistic check + generalization sweep -> diagnostics_results.json

Both scripts run in well under a minute on CPU. No downloaded weights, no API keys, no internet access required — fully self-contained and reproducible via fixed seeds.

Files

concepts.py       # synthetic activation-space generator (ConceptWorld)
adapters.py       # bias-only / scalar-affine / full-rank adapter modules
run_experiment.py # main experiment: train + eval all adapters across 4 conditions
diagnostics.py    # mechanistic check (W vs true M^-1) + training-breadth sweep
results_plot.png  # summary figure
WRITEUP.md        # full LessWrong-style report of motivation/method/results/limitations

Key result

results

Limitations

See WRITEUP.md — main ones: fully linear/synthetic toy setting, single random seed for the main table, fixed (not trainable) readout head, and the polysemantic condition isn't fully isolated from the seen/held-out split.

About

A small, fully offline research project testing a generalization failure mode in adapter-based self-interpretation methods

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages