Get started · Guides · Commands · API · Theory · Examples · Glossary
Turn T- and B-cell receptor repertoires into fixed-length numeric vectors you can cluster, visualise, and hand to a machine-learning model — at the scale of one receptor, one repertoire, or a whole cohort.
mirpy implements TCREMP: each receptor is embedded by its alignment distances to a fixed set of prototype sequences, so Euclidean distance in embedding space approximates pairwise alignment distance. That turns a set of variable-length strings into a coordinate system, which is what every downstream method actually needs.
PyPI package mirpy-lib, import name mir. Pure Python; the alignment, Pgen and sampling work is
reused from seqtree and
vdjtools.
pip install mirpy-libOptional extras, each for one job:
pip install "mirpy-lib[ann]" # pynndescent: approximate neighbours at whole-repertoire scale
pip install "mirpy-lib[ml]" # torch: the neural codecs and the learned set encoder
pip install "mirpy-lib[bench]" # the benchmark harness and theory experiments
pip install "mirpy-lib[examples]" # marimo, matplotlib, umap-learn: to run the notebooksChoosing the right scale is most of the work:
| Scale | One unit is | You get | Module |
|---|---|---|---|
| Clonotype | one receptor | a point in prototype space | mir.embedding |
| Repertoire | one sample from one donor | one fingerprint, Φ(S) |
mir.repertoire |
| Cohort | many donors, several chains each | one comparable matrix | mir.cohort |
import polars as pl
from mir.embedding.tcremp import TCREmp
model = TCREmp.from_defaults("human", "TRB") # per-chain preset: 2000 prototypes
df = pl.DataFrame({
"v_call": ["TRBV10-3*01", "TRBV20-1*01"],
"j_call": ["TRBJ2-7*01", "TRBJ1-2*01"],
"junction_aa": ["CASSIRSSYEQYF", "CSARVSGYYGYTF"],
})
X = model.embed(df) # (2, 6000) float32 — three distances per prototypeDenoise, then cluster:
from mir.embedding.pca import pca_denoise
from mir.bench.metrics import cluster
labels = cluster(pca_denoise(model.embed(vdjdb_df), n_components=65))From the shell:
mir embed clonotypes sample.tsv -o clonotypes.parquet --pca 50
mir embed repertoires cohort/*.tsv.gz -o phi.tsv --mmd mmd.tsv
mir signature --corpus blood cohort/*.tsv.gz -o rsig.parquetGetting started walks through all of this in about fifteen minutes, offline, with nothing to download.
| You want to | Go to |
|---|---|
| Embed receptors and cluster antigen-specific ones | Getting started |
| Pick a prototype count and a PCA dimension | Which prototypes |
| Find enriched or antigen-driven neighbourhoods | mir.density |
| Compare whole repertoires rather than clonotypes | mir.repertoire |
| Get one fixed, named feature vector per sample | Signatures |
| Name which part of a vector carries a signal | Channels |
| Understand why any of this is valid | Theory |
| Look up a command or a function | Commands · API |
| Check what a term means | Glossary |
| Module | What it does |
|---|---|
mir.embedding |
TCREmp and PairedTCREmp, PCA denoising, the bundled prototypes and per-chain presets |
mir.distances |
Junction distance through seqtree.gapblock, plus precomputed arda germline distances |
mir.density |
Graph-free TCRNET and ALICE neighbourhood enrichment in embedding space, with exact, kdtree and approximate backends |
mir.repertoire |
One vector per sample: kernel mean, coverage-standardised Hill diversity, second moment; MMD distance, motif witness, sub-probability measures |
mir.signature |
The portable signature — fixed, named, already-standardised columns that join to vdjtools' half on sample_id |
mir.explain |
Which named channel carries a signal under your scorer, and which clonotypes drive it |
mir.cohort |
The digital donor — multi-chain donor embeddings, batch residualisation, sample clustering |
mir.track |
Exposure trajectory — a latent progression axis disentangled from a known covariate |
mir.generate, mir.twin |
Sample new synthetic donor states, or perturb and resample one donor's state |
mir.ml |
Neural codecs, learned set encoders, diffusion generator — experimental, [ml] extra |
mir.bench |
VDJdb loader, clustering metrics, cohort scorers, reproduced theory experiments |
mir.cli |
The mir console script — embed clonotypes, embed repertoires, signature, corpus |
Every embedding is distances to prototypes, so the prototype set is the coordinate system. mirpy ships 10,000 real receptors per chain — a fixed-seed uniform sample of unique, productive, germline-resolvable clonotypes from arda-annotated repertoires, for human TRA/TRB/TRG/TRD/IGH/IGK/IGL and mouse TRA/TRB. Nothing is downloaded at import or build time.
Real rather than model-generated, because synthetic junctions have degenerate lengths and embed
measurably worse. To ask whether a result depends on which prototypes you drew, take a replicate:
the geometry agrees at R = 0.993 between two independent draws at the default n=1000, and 0.997 at
n=2000, so from about n≈500 up the draw is not what your result rests on.
Each replicate is a different coordinate system — distances within one are comparable, across two are not, and every container refuses a prototype-hash mismatch rather than mixing them silently. Details and the full presets table.
One fixed vector Φ(S) per repertoire, robust down into the low-coverage bulk-RNA-seq regime. It
sketches the repertoire as a weighted cloud of receptor points in three blocks: a random-feature
kernel mean, a coverage-standardised Hill diversity profile, and a second-moment block carrying
co-occurrence structure. Distance between two repertoires is the MMD.
from mir.repertoire import fit_repertoire_space, sample_embedding, mmd_matrix
space = fit_repertoire_space(model, pooled_clonotypes) # ONE basis for the whole cohort
embs = [sample_embedding(space, s) for s in samples]
D = mmd_matrix(embs, unbiased=True)Every sample in a cohort must go through one prototype set and one basis, or the distances are not
comparable. Use unbiased=True whenever depths differ: the biased estimator's 1/n_eff self-term
inflates low-diversity samples and manufactures signal with the wrong sign.
Φ(S) is fitted on your cohort, which is what makes it powerful and also what makes it
incomparable with anyone else's. When you need a vector a collaborator can reproduce independently,
use a signature instead: fixed, named columns standardised against a published reference corpus.
mir signature --corpus blood cohort/*.tsv.gz -o rsig.tsv # the geometry half
vdjtools signature --corpus blood cohort/*.tsv.gz -o vsig.tsv # the statistics halfThe two join on sample_id. Nine corpora are published, all at 256 components per locus; artifacts
are fetched and digest-verified on first use. --corpus is required — a signature is comparable to
another one only if both were rotated through the same corpus.
Signatures ·
Channels
CPU-parallel by default; the GPU is used only by mir.ml. Embedding runs the C++
seqtree.gapblock scorer at about 530 million pairs per second on 16 cores and releases the GIL.
Density defaults to an exact multicore kd-tree; the approximate backend is about 30× faster past 1e5
clones and is deliberately asymmetric — only the observed side is approximate, which biases
enrichment down, while the background is always exact.
Full table of parallelism settings.
Six runnable marimo notebooks — plain Python files, so they diff like source:
pip install "mirpy-lib[examples]"
marimo edit examples/signature_pipeline.py # start heresignature_pipeline (a folder of AIRR files to a joinable table — read this one first) ·
quickstart (embed, denoise, cluster, UMAP) · signature (how the halves differ, and how many
components survive a refit) · density (enrichment against a background model) ·
trajectory_and_twin (time courses and counterfactual donors) · theory (the results the
library rests on). See the
gallery.
Repo-local .venv via uv:
bash setup.sh --dev-parents --tests # editable-installs sibling seqtree / vdjtools checkouts
python -m pytest tests/ -qSee CLAUDE.md for the architecture and the reuse map.
Method: Kremlyakova et al., TCREMP: a bioinformatic pipeline for efficient embedding of T-cell receptor sequences, J Mol Biol 437 (2025) 169205.
The theory this implements is collected in the
mathematical foundations; derivations and recorded
benchmark numbers live in the companion analysis repository 2026-mirpy-analysis, with the LaTeX
appendix in 2026-mirpy-ms.
GPL-3.0-or-later. The classical v1.x/v2 repertoire toolkit — parsing, overlap, diversity, TCRnet,
GLIPH — is frozen on the legacy-v2 branch
(mirpy-lib 2.x); its functionality lives on in
vdjtools and
vdjmatch.
mir signature --corpus blood --jobs 8 samples/*.tsv -o signatures.parquet--jobs controls worker processes. It defaults to 1; use 0 to request all available
cores or a positive count to set the budget explicitly. Each spawned worker uses one
thread per numerical kernel. Workers read one sample at a time, retain their fitted
resources, and return only a feature row. Embeddings use fixed 4,096-row blocks
(24 MiB per float64 block at 256 prototypes). Output follows input order. No sample-count
heuristic changes an explicit worker budget.
For the Python cohort API, pass picklable zero-argument readers to defer file loading
into workers. Passing already-loaded frames keeps those frames in the parent too.
The CLI sets kernel thread limits before imports, including at --jobs 1. Direct
Python calls retain the calling process's Polars and BLAS settings at n_jobs=1.
Density analysis also accepts an explicit kernel budget:
neighbor_enrichment(observed, background, threads=8). Background counts are exact with both
KD-tree and ANN observation queries, including points tied at the selected radius. Only ANN
observation-neighbor recall is approximate.