biotapy.pp.filter_features#
- biotapy.pp.filter_features(adata, *, min_prevalence=None, min_total=None)#
Keep features that are present in enough samples and have enough reads.
- Parameters:
- Return type:
- Returns:
AnnData Same type as
adatawith the kept features in their original order; a TreeData keeps their subtree. Both thresholds are inclusive and, when both are given, a feature must pass both.layers,obsm,obsp,varm,varpand every non-biotapyunskey are dropped because they described the old features.- Raises:
ValueError – Neither threshold is given,
min_prevalenceis outside 0 to 1, or no feature passes.
Notes
R equivalent:
phyloseq::filter_taxaGuide: Filtering and rarefactionIn R:
filter_taxa(physeq, function(x) sum(x > 0) >= p * length(x), prune = TRUE)formin_prevalence=p, andfunction(x) sum(x) >= nformin_total=n. biotapy keeps a feature whenpresent / n_obs >= min_prevalence, while phyloseq’ssum(x > 0) >= p * length(x)can drop it at an exact boundary through floating point: 7 of 25 samples atp = 0.28, since0.28 * 25is7.000000000000001.Examples
>>> import biotapy as bt >>> bt.pp.filter_features(bt.datasets.toy(), min_prevalence=1.0).n_vars 2