biotapy.pp.filter_features

biotapy.pp.filter_features#

biotapy.pp.filter_features(adata, *, min_prevalence=None, min_total=None)#

Keep features that are present in enough samples and have enough reads.

Parameters:
  • adata (AnnData) – Samples x features.

  • min_prevalence (float | None (default: None)) – Keep features that are non-zero in at least this fraction of samples, from 0 to 1.

  • min_total (float | None (default: None)) – Keep features whose total over all samples is at least this.

Return type:

AnnData

Returns:

AnnData Same type as adata with the kept features in their original order; a TreeData keeps their subtree. Both thresholds are inclusive and, when both are given, a feature must pass both. layers, obsm, obsp, varm, varp and every non-biotapy uns key are dropped because they described the old features.

Raises:

ValueError – Neither threshold is given, min_prevalence is outside 0 to 1, or no feature passes.

Notes

R equivalent: phyloseq::filter_taxa Guide: Filtering and rarefaction

In R: filter_taxa(physeq, function(x) sum(x > 0) >= p * length(x), prune = TRUE) for min_prevalence=p, and function(x) sum(x) >= n for min_total=n. biotapy keeps a feature when present / n_obs >= min_prevalence, while phyloseq’s sum(x > 0) >= p * length(x) can drop it at an exact boundary through floating point: 7 of 25 samples at p = 0.28, since 0.28 * 25 is 7.000000000000001.

Examples

>>> import biotapy as bt
>>> bt.pp.filter_features(bt.datasets.toy(), min_prevalence=1.0).n_vars
2