biotapy.tl.permanova

Contents

biotapy.tl.permanova#

biotapy.tl.permanova(adata, grouping, *, distance='braycurtis', permutations=999, seed=None)#

Permutational multivariate analysis of variance of a distance matrix in obsp.

Parameters:
  • adata (AnnData) – Samples x features with obsp[distance], written by biotapy.tl.beta() or biotapy.tl.unifrac() with inplace=True.

  • grouping (str) – The obs column holding each sample’s group: categorical, string or bool.

  • distance (str (default: 'braycurtis')) – The obsp key to test.

  • permutations (int (default: 999)) – Permutations for the p-value.

  • seed (int | Generator | None (default: None)) – Seed or generator for the permutations.

Return type:

Series

Returns:

pandas.Series scikit-bio’s result: test statistic (pseudo-F), p-value, sample size, number of groups, number of permutations and the method and statistic names.

Raises:
  • KeyError – grouping is not an obs column, or obsp[distance] is missing.

  • TypeError – obs[grouping] is numeric (and not bool); convert it with .astype("category") for groups.

  • ValueError – grouping has missing values, or the distances hold NaN.

Notes

R equivalent: vegan::adonis2 Guide: Ordination and PERMANOVA

Matches adonis2(distance ~ grouping, data, permutations) with one term, whose F is the test statistic. P-values agree only up to permutation noise: R and NumPy random generators differ. grouping is categorical, one group per distinct value, like an R factor. adonis2 fits a numeric column as one continuous term instead, so a numeric column raises rather than silently becoming one group per value. scikit-bio’s F-statistic runs on one OpenMP thread here: with one thread per core it took 24 s instead of 0.01 s on a busy machine.

References

Anderson MJ (2001) A new method for non-parametric multivariate analysis of variance. Austral Ecology 26:32-46.

Examples

>>> import biotapy as bt
>>> tdata = bt.datasets.toy()
>>> bt.tl.beta(tdata, inplace=True)
>>> result = bt.tl.permanova(tdata, "group", seed=0)
>>> int(result["number of groups"])
2