Data model#
Every biotapy dataset is one object: a treedata TreeData
(an AnnData that also carries a phylogeny). This page covers
the one layout rule every function relies on, and what each slot on that object holds.
Samples are always rows#
X is samples (rows) by features (columns) - never the other way round. If you load data from
a tool that stores features as rows (most do), the biotapy reader transposes it for you, once,
so everything downstream can assume the same orientation.
import biotapy as bt
tdata = bt.datasets.toy()
tdata.shape # (6 samples, 8 features)
Slots#
Slot |
Holds |
Example keys |
|---|---|---|
|
Counts (or another abundance) as a sparse matrix |
- |
|
Transforms of |
|
|
Sample metadata, and per-sample results |
|
|
Taxonomy, one column per rank, and sequences |
|
|
The phylogeny, leaves named after your features |
|
|
Ordinations and embeddings |
|
|
Sample-by-sample distance matrices |
|
|
biotapy’s own bookkeeping, and ordination summaries |
|
uns["biotapy"]["x_kind"] records what X currently holds (counts, relative, rpk, cpm
or abundance); a missing key means counts. uns["biotapy"]["provenance"] lists every
biotapy function that has touched the object, in order, so you can always tell how it reached
its current state.
What survives a filter or an aggregation#
Functions that change which features exist (dropping rare taxa, aggregating to a rank) drop
layers, obsm, obsp, varm, varp, the ordination summaries uns["biotapy"]["pcoa"] and
["nmds"], and every uns key other than uns["biotapy"] - for example a plotted
uns["group_colors"] disappears along with them - because a transform, distance or annotation
computed on the old features would silently misdescribe the new ones.
Functions that only add a layer, or that subset samples, leave everything else in place.
See the data-model-slots contract for the exact rules every biotapy function follows.