Skip to content
settylabPublic

About

Differential abundance and gene expression in single-cell data

Resources

Stars

10 stars

Watchers

2 watching

Forks

Repository files navigation

Kompot

DOI PyPI Tests codecov Documentation Status

Kompot Logo

Kompot is a Python package for differential abundance and gene expression analysis using Gaussian Process models with JAX backend.

Overview

Kompot implements methodologies from the Mellon package for computing differential abundance and gene expression, with a focus on using Mahalanobis distance as a measure of differential expression significance. It leverages JAX for efficient computations and provides a scikit-learn like API with .fit() and .predict() methods.

Key features:

  • Computation of differential abundance between conditions
  • Gene expression smoothing and uncertainty estimation
  • Mahalanobis distance calculation for differential expression significance
  • JAX-accelerated computations with optional GPU support
  • Disk-backed covariance storage for sample variance estimation
  • Resource estimation and dry run for planning large analyses
  • Full scverse compatibility with direct AnnData integration
  • Visualization tools for volcano plots, heatmaps, and embeddings
  • Command-line interface for pipeline integration

Installation

pip install kompot

Or via conda:

conda install -c bioconda kompot

See the installation guide for optional dependencies and JAX GPU support.

Usage

Python API

import kompot
import anndata as ad

# Load data
adata = ad.read_h5ad("data.h5ad")

# Differential expression
kompot.de(adata, "condition", "control", "treatment")

# Differential abundance
kompot.da(adata, "condition", "control", "treatment")

# With advanced options
from kompot import GPSettings, FDRSettings

kompot.de(
    adata, "condition", "control", "treatment",
    gp=GPSettings(sigma=0.5),
    fdr=FDRSettings(threshold=0.05),
)

Sample variance: run it as a second pass

Passing sample_col turns on sample variance, which replaces the single shared posterior covariance with one (n_landmarks, n_landmarks) covariance matrix per gene, per condition. Two costs follow, and they respond to different levers:

  • Memory, 2 x n_landmarks^2 x n_genes x 8 bytes — about 0.37 GiB per gene at the default 5 000 landmarks. StorageSettings(store_arrays_on_disk=True) removes this almost entirely, and n_landmarks shrinks it quadratically.
  • Compute, one Cholesky factorisation per gene instead of one in total. store_arrays_on_disk does not help here, but lowering n_landmarks does (~0.016 s/gene at 500 against ~2.1 s at 5 000, single-threaded), as does analysing fewer genes.

So run it in two passes, and price the second one first:

# Pass 1 — all genes, no sample variance
kompot.de(adata, "condition", "Young", "Old")

mahal = "kompot_de_Young_to_Old_mahalanobis"
top_genes = adata.var.sort_values(mahal, ascending=False).head(200).index

# Pass 2 — sample variance, restricted to the top genes
plan = kompot.de(
    adata, "condition", "Young", "Old",
    sample_col="donor_id",
    genes=top_genes,
    gp=kompot.GPSettings(n_landmarks=2000),   # cost is quadratic in this
    dry_run=True,                             # drop once the plan fits
)

dry_run=True returns a full resource plan (per-array memory, disk, output fields, feasibility) without running anything; kompot de --dry-run does the same from the CLI. Details, measured plans, and the remaining levers: Planning Memory and Disk.

Command-Line Interface

# Differential expression
kompot de input.h5ad -o output.h5ad \
  --groupby condition \
  --condition1 control \
  --condition2 treatment

Documentation

Citation

If you use Kompot in your research, please cite:

@article{Otto2025.06.03.657769,
    author = {Otto, Dominik J. and Arriaga-Gomez, Erica and Thieme, Elana and Yang, Ruijin and Lee, Stanley C. and Setty, Manu},
    title = {Comparing phenotypic manifolds with Kompot: Detecting differential abundance and gene expression at single-cell resolution},
    year = {2025},
    doi = {10.1101/2025.06.03.657769},
    publisher = {Cold Spring Harbor Laboratory},
    journal = {bioRxiv},
    URL = {https://www.biorxiv.org/content/10.1101/2025.06.03.657769}
}

License

GNU General Public License v3 (GPLv3)

About

Differential abundance and gene expression in single-cell data

Resources

Stars

10 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages