Skip to content

About

HeteroSC v1

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

HeteroSC for Adapter-Based Integration and Orchestration of Complementary ScFM Experts

Data DOI Python License

Single-cell foundation models (ScFMs) such as Geneformer, scGPT, CellFM and scFoundation differ in architecture, pre-training objective, input encoding and gene vocabulary, and no single model is best everywhere. HeteroSC maps each model's representation into a shared latent space with an expert-specific adapter, and lets an input-dependent router decide how much each expert contributes for every cell, gene or cell line. New experts can be added with another adapter, without retraining any foundation model.

HeteroSC schematic

Figure: overview of HeteroSC and its three downstream tasks.

  • A · Core architecture. Four pretrained single-cell foundation models act as frozen experts (ScFM 1–4: Geneformer, scGPT, CellFM, scFoundation). Each produces a native embedding $e_i^{(k)} \in \mathbb{R}^{d_k}$ of a different size. An expert-specific adapter $f_k$ projects it into a shared latent space, giving the adapted projection $\bar h_i^{(k)} \in \mathbb{R}^{d}$. The adapted projections feed two parallel branches. Branch 1 (adaptive routing): they are concatenated ($H_i \in \mathbb{R}^{4d}$) and passed through the AdaptiveRouter, which outputs one logit per expert; a softmax turns the logits into input-dependent routing weights $\alpha_i \in \mathbb{R}^{4}$. Branch 2 (stacked representations): the same projections are stacked into $\bar H_i \in \mathbb{R}^{4\times d}$. Weighted fusion combines both branches into the fused representation $h_i = \sum_k \alpha_i^{(k)} \bar h_i^{(k)} = \alpha_i^{\top} \bar H_i \in \mathbb{R}^{d}$.
  • B · Cell type annotation. The gene expression of a single cell is encoded by the frozen experts, adapters and router into a fused cell representation, which a classification head (MLP + softmax) maps to a cell type (Pancreas, Skin, Liver, Lung and Myeloid datasets). Routing is per cell.
  • C · Gene perturbation prediction (GEARS). HeteroSC produces a fused gene embedding from the basal cell state, which takes the place of GEARS' own gene embedding. The perturbation condition is encoded separately by the GEARS encoder into a perturbation embedding. Both embeddings are combined and the GEARS decoder predicts the post-perturbation expression of every gene. Routing is per gene.
  • D · Unseen-drug response prediction. The fused cell-line representation (from the cell-line transcriptome) is concatenated with a drug representation produced by a graph convolutional network (GCN) on the molecular structure. A regression head (MLP) predicts the drug response (AUC). Routing is per cell line, and validation is leave-drug-out.

Paper: Can Heterogeneous Single-Cell Foundation Models Work Better Together? HeteroSC for Adapter-Based Integration and Orchestration of Complementary ScFM Experts. Nursyafi F. S., Hanum U. L., Fuadah Y. N., Lim K. M. (manuscript under review). Data: 10.5281/zenodo.23153537

The idea in 30 seconds

For an input $i$ (a cell, a gene or a cell line) and expert $k$ with frozen embedding $e_i^{(k)}$ (panel A):

$$\bar h_i^{(k)} = f_k\big(e_i^{(k)}\big), \qquad \alpha_i = \mathrm{softmax}\big(\mathrm{MLP}(H_i)\big),\ H_i=\big[\bar h_i^{(1)};\dots;\bar h_i^{(K)}\big], \qquad h_i = \sum_{k=1}^{K} \alpha_i^{(k)}, \bar h_i^{(k)} .$$

Only the adapters $f_k$, the router and a small task head are trained. A uniform-fusion control ($\alpha_i^{(k)}=1/K$, same adapters) separates the effect of routing from the effect of merely having several representations. Three tasks use the same core: cell type annotation (routing per cell), gene perturbation prediction with GEARS (routing per gene) and unseen-drug response prediction (routing per cell line, leave-drug-out validation). Routing weights are predictive contributions under the training objective, not measures of model quality. See docs/theory.md.

Install

git clone https://github.com/kit-cml/HeteroSC_code.git && cd HeteroSC_code
python -m venv .venv && source .venv/bin/activate            # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q                                                    # about 20 s on a CPU

Try it in two minutes (CPU, no downloads)

Open notebooks/00_quickstart_synthetic.ipynb: synthetic complementary experts, HeteroSC vs the uniform control, routing diagnostics and the zero-shot kNN probe.

from heterosc import AnnotationConfig, AnnotationTrainer, HeteroSCClassifier, make_loaders

cfg = AnnotationConfig(expert_dims={"geneformer": 768, "scgpt": 512, "cellfm": 1536, "scfoundation": 768},
                       num_classes=13)                       # fusion="adaptive" (default) or "uniform"
train, val, test = make_loaders(embeddings, labels, split, cfg.batch_size)   # embeddings: {expert: (n_cells, dim)}
model = HeteroSCClassifier(cfg)
result = AnnotationTrainer(model, cfg).fit(train, val, test, labels[split == "train"])
result["test"]["expert_utilization"]                         # mean routing weight per expert

Tutorials

Notebook What it does Needs
00_quickstart_synthetic the whole idea on synthetic data CPU
01_download_and_prepare_data download from Zenodo, verify MD5, inspect the data internet, 7-Zip/unrar
02a – 02d extract embeddings frozen embeddings from Geneformer, scGPT, scFoundation, CellFM GPU, each model's repo + checkpoint
03_cell_annotation HeteroSC vs uniform vs single experts, fine-tuned and zero-shot kNN, routing per cell type embeddings from 02
04_drug_response 10-fold leave-drug-out prediction with a GCN drug encoder embeddings from 02
05_gene_perturbation HeteroSC inside GEARS, Fusion Ladder (15 expert combinations) GEARS, gene tables
06_routing_analysis routing weights, entropy and adaptive-vs-uniform across tasks outputs of 03–05

The foundation models' checkpoints are not redistributed; get them from the official releases: Geneformer · scGPT · scFoundation · CellFM · GEARS. Each has heavy, specific dependencies, so extraction runs in a separate environment per model (templates in envs/).

Data

The record 10.5281/zenodo.23153537 contains the processed annotation datasets with their train/validation/test assignment, the drug-response tables with the leave-drug-out folds, Norman/Adamson in GEARS format and the GEARS split. Details, file lists and attribution: docs/data.md.

from heterosc.zenodo import download_data
download_data("data")          

Repository layout

src/heterosc/   modules (adapter, router, fusion) · losses · annotation · drug_response · perturbation
                zenodo / gdsc (data download) · extraction/ (4 foundation models) · plots · synthetic
notebooks/      00–06 tutorials          scripts/   command-line versions
docs/           theory · data · reproducibility · method-to-code map          tests/   unit tests (run in CI)
assets/         schematic.png (figure above)

Status

Component State
Core modules, losses, annotation, kNN probe, drug-response training (with leave-drug-out guard), gene-level fusion module, Zenodo download, notebook 00 unit-tested and run on CPU (pytest)
Foundation-model extraction (heterosc.extraction), GEARS integration notebook (05) ported from working research notebooks; need each model's code, checkpoint and a GPU; not exercised by the tests
Notebooks 01–04, 06 written against the tested API; run them on your data and report problems as issues

Default training settings are in docs/reproducibility.md. You retrain everything from the foundation-model embeddings, so exact numbers will vary with seeds, hardware and library versions.

Citation

If you use this code or data, please cite the manuscript (under review) and the data record:

Nursyafi FS, Hanum UL, Fuadah YN, Lim KM. HeteroSC benchmark data: processed single-cell annotation datasets, drug-response tables, perturbation splits and data partitions. Zenodo. https://doi.org/10.5281/zenodo.23153537

See CITATION.cff. Third-party data keep the licences and terms of their original providers (see docs/data.md).

License

MIT (see LICENSE) for the code in this repository. Pre-trained models and datasets keep their own licences.

About

HeteroSC v1

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages