biotapy.io.read_metaphlan

Contents

biotapy.io.read_metaphlan#

biotapy.io.read_metaphlan(path)#

Read a MetaPhlAn profile, or several merged, into a samples x clades table.

Parameters:

path (str | Path) – A MetaPhlAn 3 or 4 profile (-t rel_ab, the default, or -t rel_ab_w_read_stats), or a table of several merged by merge_metaphlan_tables.py or with one column per sample. Gzip (.gz) is read directly.

Return type:

TreeData

Returns:

TreeData Relative abundances in X (MetaPhlAn’s percentages divided by 100), one feature per leaf clade: the deepest row of each lineage, such as MetaPhlAn 4’s SGBs (t__SGB1871) or MetaPhlAn 3’s species. var_names are the leaf’s last name without its rank prefix (SGB1871, Bacteroides_ovatus); var holds the rank columns kingdom to species. UNCLASSIFIED (UNKNOWN in older tables) stays a feature with every rank NaN, so each sample sums to 1. A single profile’s sample is named after its file, and _profile is removed from sample names, as merge_metaphlan_tables.py does. There is no tree.

Raises:

ValueError – The file is empty, not UTF-8 text, or a .gz that is not valid gzip; the header repeats a column name or has an empty sample column name; a data row has more cells than the header, or no (or a blank) id; an abundance is missing (a short row or an empty cell), not a number, negative or not finite; the leaf clades of a sample do not sum to 100 (rows removed, or a table that is not a profile, such as one with ; lineages); or leaf names or sample names repeat. Messages name path.

Notes

R equivalent: mia::importMetaPhlAn Guide: Reading and writing data

A profile lists every rank, and a clade’s abundance is the sum of its children’s, so keeping only the leaves keeps all the abundance once. bt.pp.tax_glom gives back the higher ranks; it drops UNCLASSIFIED unless dropna=False. uns['biotapy']['x_kind'] is "relative". NCBI taxids, additional_species, coverage and read estimates are not read.

The reader builds one dense clades x samples float64 array of every row the file prints before keeping the leaves as CSR (about 12 MB for HMP2’s 932-row x 1,638-sample table).

References

Blanco-Míguez A et al. (2023) Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nature Biotechnology 41:1633-1644.

Examples

>>> import tempfile
>>> from pathlib import Path
>>> import biotapy as bt
>>> rows = ["k__Bacteria\t2\t90.0\t", "k__Bacteria|g__Bacteroides\t2|816\t90.0\t", "UNCLASSIFIED\t-1\t10.0\t"]
>>> text = "#mpa_vJan25\n#clade_name\tNCBI_tax_id\trelative_abundance\tadditional_species\n" + "\n".join(rows)
>>> with tempfile.TemporaryDirectory() as tmp:
...     path = Path(tmp) / "S1_profile.tsv"
...     _ = path.write_text(text)
...     tdata = bt.io.read_metaphlan(path)
>>> tdata.obs_names.tolist(), tdata.var_names.tolist(), tdata.X.toarray().tolist()
(['S1'], ['Bacteroides', 'UNCLASSIFIED'], [[0.9, 0.1]])