Installation
pip install piaso-toolsThe distribution is piaso-tools; the import is import piaso. Pre-compiled
wheels ship for Linux, macOS (Intel and Apple Silicon) and Windows on Python
3.9–3.12, so there is nothing to build.
This also installs cytome, the on-disk dataset format PIASO reads and writes, and COSG for marker genes.
From bioconda
conda install -c conda-forge -c bioconda piasoThe bioconda package is currently 1.0.3, an older release that predates
cytome support, the exact ligand–receptor null and the rest of the 1.2 series,
and it installs scanpy as a dependency. Until the recipe catches up, pip install piaso-tools is the supported route — it works inside a conda
environment too.
Development version
pip install git+https://github.com/genecell/PIASO.gitOptional extras
pip install 'piaso-tools[scanpy]'scanpy is only needed for interoperability with scanpy-based workflows. The core workflow — normalization, dimensionality reduction, neighbours, clustering, UMAP, marker genes and plotting — runs without it.
On a cluster
Three things decide the wall time on a shared cluster, and none of them is a setting in PIASO.
Cores. The streaming engines use every core the job is allowed
(--cpus-per-task for sbatch/srun, or the core count in an Open OnDemand
form) with no argument; n_threads only caps that. More threads than cores
just take turns.
Storage. A network file system drops its cache of a file whenever the
file may have changed: on every file lock, and whenever the file is written.
A cytome is one file, read many times and written to between the reads, so
from a network mount every write (a quantified matrix, a metadata entry, a
checkpoint) is followed by a re-read of the store on the next pass. Measured
on one node, same twenty cores, a 31k-cell store from a network mount against
node-local scratch: PICCO 50.8 s against 12.8 s, the peak quantifier 41.5 s
against 15.5 s, COSG over five million tiles 2 min 58 s against 1 min 51 s.
runSVD reads without locks and is within a factor of 1.5. Copy the store to
node-local scratch before opening it, and copy it back if you write to it:
cp /path/to/data.cytome* "$TMPDIR"/ # the .cytome and its -wal/-shm sidecarspython -c "import cytome; ds = cytome.open('$TMPDIR/data.cytome')"cytome.open says so once when it sees a store on a network file system.
When a run is still slow, two lines say why:
python -c "import os; print(len(os.sched_getaffinity(0)))" # the cores you actually havePIASO_SVD_PROFILE=1 python your_script.py # per pass: threads, reading vs computingThe profile goes to standard error, which in a notebook is the server’s log rather than the cell; run the script from a terminal to see it.
An environment for the tutorials
The tutorials run on a plain PIASO install — there is no separate tutorial dependency set:
conda create -n piaso_env python=3.10 -yconda activate piaso_envpip install piaso-tools
python -c "import piaso; print(piaso.__version__)"To run them as notebooks, register the environment as a Jupyter kernel:
pip install ipykernelpython -m ipykernel install --user --name piaso_env --display-name "Python (piaso_env)"Two packages that similar guides ask for are not needed here:
| often suggested | why not |
|---|---|
pip install 'scanpy[leiden]' | piaso.tl.leiden is PIASO’s own Leiden, in its Rust extension, with igraph (installed with PIASO) as the other backend. PIASO has no scanpy dependency; pip install 'piaso-tools[scanpy]' exists only for interoperating with existing scanpy code. |
pip install scrublet | piaso.pp.scrublet is PIASO’s own implementation of the algorithm, streaming and library-aware. |
cytome comes with piaso-tools; the tutorials that stream from disk need
nothing extra.