biotapy.fn.load_hierarchy

Contents

biotapy.fn.load_hierarchy#

biotapy.fn.load_hierarchy(path, level, *, layout='parent_first')#

Read a function mapping file into an edge table for func_glom.

Parameters:
  • path (str | Path) – A tab-separated mapping file with no header line, gzip (.gz) allowed, UTF-8 with or without a BOM. Surrounding whitespace is stripped from every id (humann_regroup_table strips only each line’s ends). Blank lines and lines starting with # are skipped (humann_regroup_table does not skip # lines; it would read them as ids). Each line holds one id and then one or more ids it maps to or from (see layout); a line may repeat its first id.

  • level (str) – Name for the level the parents form, for example "pathway"; func_glom selects edges by it.

  • layout (Literal['parent_first', 'child_first'] (default: 'parent_first')) – "parent_first": parent<TAB>child<TAB>child..., the format of humann_regroup_table --custom and PICRUSt2 mapping files. "child_first": child<TAB>parent<TAB>parent..., the format of humann_regroup_table --reversed and of two-column child, parent tables.

Return type:

DataFrame

Returns:

pandas.DataFrame Columns child, parent, level and parent_name (NaN: mapping files name no parents), one row per distinct pair. attrs["source"] is the file’s path. There is no attrs["license"] (bt.datasets.enzyme sets one): biotapy cannot know your file’s licence, so set it yourself if func_glom’s provenance should record it.

Raises:

ValueError – layout is not one of the two names; the file holds no edges; a line holds a single id; or a line has an empty cell before its last id (for example a leading tab), which would make a wrong edge.

Notes

R equivalent: none Guide: Function

biotapy ships no KEGG or MetaCyc mapping: their licences forbid redistributing them. Point path at files you are licensed to use. For EC numbers, bt.datasets.enzyme() gives the open ENZYME hierarchy.

Examples

>>> import tempfile
>>> from pathlib import Path
>>> import biotapy as bt
>>> path = Path(tempfile.mkdtemp()) / "map.tsv"
>>> _ = path.write_text("P1\tK1\tK2\nP2\tK2\n")
>>> bt.fn.load_hierarchy(path, "pathway")[["child", "parent"]].values.tolist()
[['K1', 'P1'], ['K2', 'P1'], ['K2', 'P2']]