Contributing guide#

All contributions follow rules.md, the binding development rules for every change, human or agent. Background and reasoning for those rules live in the knowledge bundle (decisions, contracts, roadmap, playbooks); new public functions follow the add-a-function playbook.

We assume you are already familiar with git and with making pull requests on GitHub. For the absolute basics, see the pyopensci tutorials or the scientific Python tutorials.

Setting up a development environment#

biotapy uses uv to manage its environment. Sync the groups you need:

uv sync --group dev --group test --group doc
  • dev: ruff, mypy, import-linter, prek, asv - linting, type-checking and benchmarks.

  • test: pytest, hypothesis, coverage.

  • doc: sphinx, myst-nb, sphinx-book-theme and the other packages that build this site.

The .venv directory uv sync creates is typically auto-discovered by IDEs such as VS Code.

Linting and type-checking#

uvx prek run --all-files

This runs the same hooks as CI’s lint job: ruff lint and format (including the size/complexity limits in rules.md R5), mypy --strict on src/biotapy and docs/extensions, import-linter (the module-layer contract), and pyproject-fmt.

Running tests#

uv run --group test pytest

Network and golden tests are excluded by default ([tool.pytest] in pyproject.toml sets -m "not network and not r"). Run them explicitly:

uv run --group test pytest -m "network or golden"

These tests download bt.datasets.global_patterns(), bt.datasets.enterotype() and bt.datasets.esophagus() through pooch and compare biotapy with R on that data: pp.relative, pp.tax_glom, filtering, rarefaction (its invariants), alpha and beta diversity, UniFrac, PCoA, NMDS and PERMANOVA against phyloseq, vegan, ape and picante. Set BIOTAPY_DATA_DIR to point the pooch cache somewhere other than the default per-user cache directory - CI caches it across runs the same way:

BIOTAPY_DATA_DIR=.pooch uv run --group test pytest -m "network or golden"

Regenerating the R golden files#

The golden CSVs under tests/golden/ and the R-written fixtures under tests/data/phyloseq/ and tests/data/dada2/ are produced by a pinned R container, not by pytest, and are never hand-edited. See the regenerate-golden-files playbook for the steps and its bit-identical check.

Building the docs locally#

uv run --group doc sphinx-build -W -b html docs docs/_build/html

Then open docs/_build/html/index.html. Read the Docs builds this same site with uvx hatch run docs:build (see .readthedocs.yaml and the docs hatch environment in pyproject.toml), which wraps the equivalent sphinx-build invocation.

Notebooks run on every build (nb_execution_mode = "cache"), but a notebook is re-executed only when its own content changes. After a code change, run uvx hatch run docs:clean first, which deletes every git-ignored file under docs/ (_build, with the jupyter cache in it, and generated), or delete docs/_build/.jupyter_cache, to force re-execution locally. CI and Read the Docs always start clean.

If you refer to objects from another package, add an entry to intersphinx_mapping in docs/conf.py so Sphinx can link to it. If the build fails over a link outside your control, add an exception to nitpick_ignore in the same file.

Commit conventions#

Commits follow Conventional Commits (feat:, fix:, refactor:, docs:, test:, build:, ci:, chore:), one logical change per commit (rules.md R13.1).

Publishing a release#

Releases follow the cut-a-release playbook: bump the version, move the [Unreleased] changelog entry, tag, publish a GitHub release, and let release.yaml upload to PyPI through trusted publishing.