Every tutorial on this page is executed against the real dataset it names, on
the machine that builds this site — the numbers and the figures are what the
code produced, not what it should produce. Most fetch their data through
piaso.data, which downloads and caches it for you; a few use datasets hosted
by their original providers and give the download command on the page.
scRNA-seq
Methods
| what it covers |
|---|
| GDR | Marker-gene-guided dimensionality reduction — an embedding built on cell identity rather than variance. |
| Cell type prediction with GDR | Transferring labels from an annotated reference. |
| projectGDR | Freeze a reference’s GDR space and project new data into it — the reference’s coordinates stay fixed. |
| Leiden at scale | 200,000 cells clustered in three seconds, with the same labels on any number of threads. |
| LeidenLocal | Re-clustering inside one cluster, without re-running the whole analysis. |
| GDR at scale | 200,000 cells in 17 minutes, streamed — what the embedding costs at that size. |
| GDR and SVD on 1.5 million cells | Both embeddings over one human cortex atlas: runtime, peak memory, and whether either separates cell types better. |
| GDR on developmental data | Groups that are stages rather than terminal cell types: 3× better separation for 1.6% accuracy. |
| GDR beyond transcriptomics | Images as expression matrices — what GDR separates on data that is not single-cell at all. |
| Emergene | Condition-specific gene programs scored per cell rather than per cluster, so heterogeneous responses stay visible. |
Marker genes
| what it covers |
|---|
| COSG: marker genes and significance | Cosine-specificity markers, the p-value columns added in v1.2.0, IQR normalisation for comparing across cell types, and the one way of using the p-values that is wrong. |
| COSG on a cytome | The same markers streamed from disk, and the four output shapes the streaming path returns. |
| COSG across batches | batch_key scores each batch separately and averages — 86% of markers hold, and the 14% that move name your most dissociation-sensitive cell types. |
| COSG on the GPU | One argument, measured across five matrix sizes: 2× in the useful range, slower below 10,000 cells. |
| COSG on spatial data | Organ markers on a whole embryo section, plotted back into tissue space — the check dissociated data cannot give you. |
Annotation
| what it covers |
|---|
| Marker-based cell type prediction | Label cells from a reference’s markers, checked against 8,738 held-out cells across 20 types (95.4%) and against PIASOmarkerDB. |
| PIASOmarkerDB API client | Query 36 curated marker studies from Python, and go from a gene list back to the cell types that match. |
Gene sets and sequence
| what it covers |
|---|
| Gene set scoring (PIASOscore) | Scoring a gene set — or a whole pathway database — against matched control sets, with per-cell p-values. |
| KEGG and drug-target gene sets | 320 pathways and 659 drug target sets scored per cell, then asked which cell type each belongs to. |
| Motif analysis | Genome → promoters → PWM scan → enrichment, and the background choice that decides the answer. |
Spatial transcriptomics
| what it covers |
|---|
| Xenium with tissue-image overlay | Clusters drawn over the morphology image, and selecting a region of interest. |
| Xenium into a cytome directly | Straight from the platform output to a cytome, with no AnnData in between. |
| Xenium Prime 5K: mouse brain end to end | 63,173 cells × 5,006 genes from the raw bundle to annotated clusters, checked against known anatomy — and how to pull 83 MB out of a 13 GB archive. |
| Atera WTA | 18,028 targets in situ — 90% of the protein-coding transcriptome — and what that finds which a 313-gene panel cannot. |
| Downstream in situ | KEGG pathways per cell, LARIS ligand–receptor and cytorete regulons on both sections — and the coverage measurement that decides which are possible on your panel. |
| MERFISH sections | Several sections in one cytome, aligned and analysed together. |
| Stereo-seq whole embryo | A whole-embryo section at bin resolution, including rotating the coordinates. |
| GDR on spatial transcriptomics | Eight embryonic stages, 520,815 bins, in one embedding. |
Gene regulatory networks
Cell-cell interaction
| what it covers |
|---|
| SCALAR | Ligand-receptor pairs between cell types, with a KNN-matched permutation null. 1.7M interactions in under a minute. |
| LARIS | The spatial counterpart: when cells have coordinates, proximity constrains which interactions are possible. Links the six worked tutorials that ship with the package. |
Plotting and data
| what it covers |
|---|
| Plotting | piaso.pl end to end: embeddings, dot plots, violins, splits. |
| Colour palettes | The built-in palettes, and how to set your own. |
| Datasets and genome references | piaso.data: what is available, and how caching works. |
| cytome basics | The file format itself — what it stores and how to read it. |
| Converting: AnnData, Seurat, SingleCellExperiment | cytome as an interchange format, and what does and does not travel with it. |
| cytome in R | Read, write and stream .cytome natively in R, with no Python runtime. |
| Agents and project tooling | PIASO-for-agents, stato and PlanDrop — running long analyses with coding agents. |
Which one first?
Start with
Human PBMC scRNA-seq end to end (AnnData)
if your data is human, and
Mouse brain scRNA-seq end to end (AnnData)
if it is mouse. Either introduces every function the others reuse; the pair
differs in what the QC step has to do, which is the part that does not transfer
between samples. Median mitochondrial content is 12.0% in the human PBMC
sample and 0.011% in the mouse nuclei, so on the human page every threshold
bites and one cluster turns out to be mitochondrial reads, while on the mouse
page the same thresholds are inert and the highest-mitochondrial cluster is
endothelium that should be kept.
Then:
- more cells than fit in memory → the cytome version of the same page;
- more than one library → multiple samples;
- clusters in hand and markers wanted →
COSG.
Each AnnData page has a cytome twin running deliberately the same analysis.
Reading a pair side by side is the fastest way to see what changes when the
matrix stays on disk: almost nothing in the calls, everything in where the
results live.
From the previous release
Thirteen more tutorials are published under /tutorials/previous/. They
are listed here rather than only counted, because several cover ground no
current page does — CellRanger pre-processing, and GDR on scATAC-seq.
They were recovered from the PIASO v1.1.0 documentation and have not been
re-run against this release, so each carries a banner saying so. They are
worth reading for the analysis; check any call against the
API reference before relying on it.
The original v1.1.0 site remains available at
https://genecell.github.io/PIASO/ if you need the pages exactly as they
were published, with their original outputs.