Kompot is a Python package for differential abundance and gene expression analysis using Gaussian Process models with JAX backend.
Kompot implements methodologies from the Mellon package for computing differential abundance and gene expression, with a focus on using Mahalanobis distance as a measure of differential expression significance. It leverages JAX for efficient computations and provides a scikit-learn like API with .fit() and .predict() methods.
Key features:
- Computation of differential abundance between conditions
- Gene expression smoothing and uncertainty estimation
- Mahalanobis distance calculation for differential expression significance
- JAX-accelerated computations with optional GPU support
- Disk-backed covariance storage for sample variance estimation
- Resource estimation and dry run for planning large analyses
- Full scverse compatibility with direct AnnData integration
- Visualization tools for volcano plots, heatmaps, and embeddings
- Command-line interface for pipeline integration
pip install kompotOr via conda:
conda install -c bioconda kompotSee the installation guide for optional dependencies and JAX GPU support.
import kompot
import anndata as ad
# Load data
adata = ad.read_h5ad("data.h5ad")
# Differential expression
kompot.de(adata, "condition", "control", "treatment")
# Differential abundance
kompot.da(adata, "condition", "control", "treatment")
# With advanced options
from kompot import GPSettings, FDRSettings
kompot.de(
adata, "condition", "control", "treatment",
gp=GPSettings(sigma=0.5),
fdr=FDRSettings(threshold=0.05),
)Passing sample_col turns on sample variance, which replaces the single
shared posterior covariance with one (n_landmarks, n_landmarks) covariance
matrix per gene, per condition. Two costs follow, and they respond to
different levers:
- Memory,
2 x n_landmarks^2 x n_genes x 8bytes — about 0.37 GiB per gene at the default 5 000 landmarks.StorageSettings(store_arrays_on_disk=True)removes this almost entirely, andn_landmarksshrinks it quadratically. - Compute, one Cholesky factorisation per gene instead of one in total.
store_arrays_on_diskdoes not help here, but loweringn_landmarksdoes (~0.016 s/gene at 500 against ~2.1 s at 5 000, single-threaded), as does analysing fewer genes.
So run it in two passes, and price the second one first:
# Pass 1 — all genes, no sample variance
kompot.de(adata, "condition", "Young", "Old")
mahal = "kompot_de_Young_to_Old_mahalanobis"
top_genes = adata.var.sort_values(mahal, ascending=False).head(200).index
# Pass 2 — sample variance, restricted to the top genes
plan = kompot.de(
adata, "condition", "Young", "Old",
sample_col="donor_id",
genes=top_genes,
gp=kompot.GPSettings(n_landmarks=2000), # cost is quadratic in this
dry_run=True, # drop once the plan fits
)dry_run=True returns a full resource plan (per-array memory, disk, output
fields, feasibility) without running anything; kompot de --dry-run does the
same from the CLI. Details, measured plans, and the remaining levers:
Planning Memory and Disk.
# Differential expression
kompot de input.h5ad -o output.h5ad \
--groupby condition \
--condition1 control \
--condition2 treatment- Full Documentation
- Planning Memory and Disk — what sample variance costs and the two-pass workflow
- Tutorial Notebooks
- Getting Started — differential expression, end to end
- Advanced Differential Expression — tuning, multiple comparisons, run tracking, resource planning
- DE with Sample Variance — replicate-aware significance
- Differential Abundance — cell-state frequency changes, incl. sample variance
- Smoothing Expression — the expression function underneath DE: its two uncertainties, and fit/predict across cells
- CLI Guide
If you use Kompot in your research, please cite:
@article{Otto2025.06.03.657769,
author = {Otto, Dominik J. and Arriaga-Gomez, Erica and Thieme, Elana and Yang, Ruijin and Lee, Stanley C. and Setty, Manu},
title = {Comparing phenotypic manifolds with Kompot: Detecting differential abundance and gene expression at single-cell resolution},
year = {2025},
doi = {10.1101/2025.06.03.657769},
publisher = {Cold Spring Harbor Laboratory},
journal = {bioRxiv},
URL = {https://www.biorxiv.org/content/10.1101/2025.06.03.657769}
}GNU General Public License v3 (GPLv3)
