Single-cell foundation models (ScFMs) such as Geneformer, scGPT, CellFM and scFoundation differ in architecture, pre-training objective, input encoding and gene vocabulary, and no single model is best everywhere. HeteroSC maps each model's representation into a shared latent space with an expert-specific adapter, and lets an input-dependent router decide how much each expert contributes for every cell, gene or cell line. New experts can be added with another adapter, without retraining any foundation model.
Figure: overview of HeteroSC and its three downstream tasks.
-
A · Core architecture. Four pretrained single-cell foundation models act as frozen experts (ScFM 1–4: Geneformer, scGPT, CellFM, scFoundation). Each produces a native embedding
$e_i^{(k)} \in \mathbb{R}^{d_k}$ of a different size. An expert-specific adapter$f_k$ projects it into a shared latent space, giving the adapted projection$\bar h_i^{(k)} \in \mathbb{R}^{d}$ . The adapted projections feed two parallel branches. Branch 1 (adaptive routing): they are concatenated ($H_i \in \mathbb{R}^{4d}$ ) and passed through the AdaptiveRouter, which outputs one logit per expert; a softmax turns the logits into input-dependent routing weights$\alpha_i \in \mathbb{R}^{4}$ . Branch 2 (stacked representations): the same projections are stacked into$\bar H_i \in \mathbb{R}^{4\times d}$ . Weighted fusion combines both branches into the fused representation$h_i = \sum_k \alpha_i^{(k)} \bar h_i^{(k)} = \alpha_i^{\top} \bar H_i \in \mathbb{R}^{d}$ . - B · Cell type annotation. The gene expression of a single cell is encoded by the frozen experts, adapters and router into a fused cell representation, which a classification head (MLP + softmax) maps to a cell type (Pancreas, Skin, Liver, Lung and Myeloid datasets). Routing is per cell.
- C · Gene perturbation prediction (GEARS). HeteroSC produces a fused gene embedding from the basal cell state, which takes the place of GEARS' own gene embedding. The perturbation condition is encoded separately by the GEARS encoder into a perturbation embedding. Both embeddings are combined and the GEARS decoder predicts the post-perturbation expression of every gene. Routing is per gene.
- D · Unseen-drug response prediction. The fused cell-line representation (from the cell-line transcriptome) is concatenated with a drug representation produced by a graph convolutional network (GCN) on the molecular structure. A regression head (MLP) predicts the drug response (AUC). Routing is per cell line, and validation is leave-drug-out.
Paper: Can Heterogeneous Single-Cell Foundation Models Work Better Together? HeteroSC for Adapter-Based Integration and Orchestration of Complementary ScFM Experts. Nursyafi F. S., Hanum U. L., Fuadah Y. N., Lim K. M. (manuscript under review). Data: 10.5281/zenodo.23153537
For an input
Only the adapters docs/theory.md.
git clone https://github.com/kit-cml/HeteroSC_code.git && cd HeteroSC_code
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q # about 20 s on a CPUOpen notebooks/00_quickstart_synthetic.ipynb: synthetic complementary experts, HeteroSC vs the uniform control, routing diagnostics and the zero-shot kNN probe.
from heterosc import AnnotationConfig, AnnotationTrainer, HeteroSCClassifier, make_loaders
cfg = AnnotationConfig(expert_dims={"geneformer": 768, "scgpt": 512, "cellfm": 1536, "scfoundation": 768},
num_classes=13) # fusion="adaptive" (default) or "uniform"
train, val, test = make_loaders(embeddings, labels, split, cfg.batch_size) # embeddings: {expert: (n_cells, dim)}
model = HeteroSCClassifier(cfg)
result = AnnotationTrainer(model, cfg).fit(train, val, test, labels[split == "train"])
result["test"]["expert_utilization"] # mean routing weight per expert| Notebook | What it does | Needs |
|---|---|---|
00_quickstart_synthetic |
the whole idea on synthetic data | CPU |
01_download_and_prepare_data |
download from Zenodo, verify MD5, inspect the data | internet, 7-Zip/unrar |
02a – 02d extract embeddings |
frozen embeddings from Geneformer, scGPT, scFoundation, CellFM | GPU, each model's repo + checkpoint |
03_cell_annotation |
HeteroSC vs uniform vs single experts, fine-tuned and zero-shot kNN, routing per cell type | embeddings from 02 |
04_drug_response |
10-fold leave-drug-out prediction with a GCN drug encoder | embeddings from 02 |
05_gene_perturbation |
HeteroSC inside GEARS, Fusion Ladder (15 expert combinations) | GEARS, gene tables |
06_routing_analysis |
routing weights, entropy and adaptive-vs-uniform across tasks | outputs of 03–05 |
The foundation models' checkpoints are not redistributed; get them from the official releases:
Geneformer · scGPT · scFoundation · CellFM · GEARS.
Each has heavy, specific dependencies, so extraction runs in a separate environment per model (templates in envs/).
The record 10.5281/zenodo.23153537 contains the processed annotation datasets with their train/validation/test assignment, the drug-response tables with the leave-drug-out folds,
Norman/Adamson in GEARS format and the GEARS split. Details, file lists and attribution: docs/data.md.
from heterosc.zenodo import download_data
download_data("data") src/heterosc/ modules (adapter, router, fusion) · losses · annotation · drug_response · perturbation
zenodo / gdsc (data download) · extraction/ (4 foundation models) · plots · synthetic
notebooks/ 00–06 tutorials scripts/ command-line versions
docs/ theory · data · reproducibility · method-to-code map tests/ unit tests (run in CI)
assets/ schematic.png (figure above)
| Component | State |
|---|---|
| Core modules, losses, annotation, kNN probe, drug-response training (with leave-drug-out guard), gene-level fusion module, Zenodo download, notebook 00 | unit-tested and run on CPU (pytest) |
Foundation-model extraction (heterosc.extraction), GEARS integration notebook (05) |
ported from working research notebooks; need each model's code, checkpoint and a GPU; not exercised by the tests |
| Notebooks 01–04, 06 | written against the tested API; run them on your data and report problems as issues |
Default training settings are in docs/reproducibility.md. You retrain everything from the foundation-model embeddings, so exact numbers will vary with seeds, hardware and library versions.
If you use this code or data, please cite the manuscript (under review) and the data record:
Nursyafi FS, Hanum UL, Fuadah YN, Lim KM. HeteroSC benchmark data: processed single-cell annotation datasets, drug-response tables, perturbation splits and data partitions. Zenodo. https://doi.org/10.5281/zenodo.23153537
See CITATION.cff. Third-party data keep the licences and terms of their original providers (see docs/data.md).
MIT (see LICENSE) for the code in this repository. Pre-trained models and datasets keep their own licences.
