Skip to content

Installation

Terminal window
pip install piaso-tools

The distribution is piaso-tools; the import is import piaso. Pre-compiled wheels ship for Linux, macOS (Intel and Apple Silicon) and Windows on Python 3.9–3.12, so there is nothing to build.

This also installs cytome, the on-disk dataset format PIASO reads and writes, and COSG for marker genes.

From bioconda

Terminal window
conda install -c conda-forge -c bioconda piaso

The bioconda package is currently 1.0.3, an older release that predates cytome support, the exact ligand–receptor null and the rest of the 1.2 series, and it installs scanpy as a dependency. Until the recipe catches up, pip install piaso-tools is the supported route — it works inside a conda environment too.

Development version

Terminal window
pip install git+https://github.com/genecell/PIASO.git

Optional extras

Terminal window
pip install 'piaso-tools[scanpy]'

scanpy is only needed for interoperability with scanpy-based workflows. The core workflow — normalization, dimensionality reduction, neighbours, clustering, UMAP, marker genes and plotting — runs without it.

On a cluster

Three things decide the wall time on a shared cluster, and none of them is a setting in PIASO.

Cores. The streaming engines use every core the job is allowed (--cpus-per-task for sbatch/srun, or the core count in an Open OnDemand form) with no argument; n_threads only caps that. More threads than cores just take turns.

Storage. A network file system drops its cache of a file whenever the file may have changed: on every file lock, and whenever the file is written. A cytome is one file, read many times and written to between the reads, so from a network mount every write (a quantified matrix, a metadata entry, a checkpoint) is followed by a re-read of the store on the next pass. Measured on one node, same twenty cores, a 31k-cell store from a network mount against node-local scratch: PICCO 50.8 s against 12.8 s, the peak quantifier 41.5 s against 15.5 s, COSG over five million tiles 2 min 58 s against 1 min 51 s. runSVD reads without locks and is within a factor of 1.5. Copy the store to node-local scratch before opening it, and copy it back if you write to it:

Terminal window
cp /path/to/data.cytome* "$TMPDIR"/ # the .cytome and its -wal/-shm sidecars
python -c "import cytome; ds = cytome.open('$TMPDIR/data.cytome')"

cytome.open says so once when it sees a store on a network file system.

When a run is still slow, two lines say why:

Terminal window
python -c "import os; print(len(os.sched_getaffinity(0)))" # the cores you actually have
PIASO_SVD_PROFILE=1 python your_script.py # per pass: threads, reading vs computing

The profile goes to standard error, which in a notebook is the server’s log rather than the cell; run the script from a terminal to see it.

An environment for the tutorials

The tutorials run on a plain PIASO install — there is no separate tutorial dependency set:

Terminal window
conda create -n piaso_env python=3.10 -y
conda activate piaso_env
pip install piaso-tools
python -c "import piaso; print(piaso.__version__)"

To run them as notebooks, register the environment as a Jupyter kernel:

Terminal window
pip install ipykernel
python -m ipykernel install --user --name piaso_env --display-name "Python (piaso_env)"

Two packages that similar guides ask for are not needed here:

often suggestedwhy not
pip install 'scanpy[leiden]'piaso.tl.leiden is PIASO’s own Leiden, in its Rust extension, with igraph (installed with PIASO) as the other backend. PIASO has no scanpy dependency; pip install 'piaso-tools[scanpy]' exists only for interoperating with existing scanpy code.
pip install scrubletpiaso.pp.scrublet is PIASO’s own implementation of the algorithm, streaming and library-aware.

cytome comes with piaso-tools; the tutorials that stream from disk need nothing extra.