biotapy.io.read_biom

Contents

biotapy.io.read_biom#

biotapy.io.read_biom(path, *, tree=None)#

Read a BIOM table (JSON 1.0 or HDF5 2.1) with samples as rows.

Parameters:
  • path (str | Path) – BIOM file; JSON or HDF5 is detected from the content.

  • tree (str | Path | None (default: None)) – Newick file whose tips are the table’s observation ids.

Return type:

TreeData

Returns:

TreeData The table in X; observation taxonomy metadata as rank columns in var; sample metadata in obs, with empty values as NaN; the tree in vart['phylo'].

Raises:
  • biom.exception.TableException – The file repeats a sample or observation id (raised by biom-format).

  • ValueError – tree is not a valid Newick tree, or shares no tip with the table.

Warns:

UserWarning – Tree tips and table features differ; only shared features are kept.

Notes

R equivalent: phyloseq::import_biom Guide: Reading and writing data

BIOM does not record what X holds, so uns['biotapy']['x_kind'] is inferred from the values: non-negative whole numbers are "counts", rows that each sum to 1 are "relative", anything else is "abundance".

Only the taxonomy (or Taxonomy) observation metadata key is read; other observation keys, such as confidence, are not.

Examples

>>> import tempfile
>>> from pathlib import Path
>>> import biom
>>> from biom.util import biom_open
>>> import biotapy as bt
>>> toy = bt.datasets.toy()
>>> path = Path(tempfile.mkdtemp()) / "toy.biom"
>>> table = biom.Table(toy.X.T, list(toy.var_names), list(toy.obs_names))
>>> with biom_open(str(path), "w") as handle:
...     table.to_hdf5(handle, "example")
>>> bt.io.read_biom(path).shape
(6, 8)