Skip to content

Tutorials

Every tutorial on this page is executed against the real dataset it names, on the machine that builds this site — the numbers and the figures are what the code produced, not what it should produce. Most fetch their data through piaso.data, which downloads and caches it for you; a few use datasets hosted by their original providers and give the download command on the page.

scRNA-seq

what it covers
Human PBMC scRNA-seq end to end (AnnData)Counts → two-tailed QC → clusters → cell types, on human blood. A whole cluster turns out to be mitochondrial reads. The place to start.
Human PBMC scRNA-seq end to end (cytome)The same human analysis streamed from disk, so memory does not scale with cell count.
Mouse brain scRNA-seq end to end (AnnData)Cell Ranger counts → QC → doublets → clusters → cell types, on mouse brain nuclei.
Mouse brain scRNA-seq end to end (cytome)The same mouse analysis, streamed.
Multiple samplesPer-library doublets, per-sample QC, and what a batch effect does and does not look like.
Cells vs nucleiThe same tissue prepared both ways — a real batch effect, measured, and where it comes from.
Multiple samples in one cytomeThe same multi-sample analysis with every library streamed from a single file.

Methods

what it covers
GDRMarker-gene-guided dimensionality reduction — an embedding built on cell identity rather than variance.
Cell type prediction with GDRTransferring labels from an annotated reference.
projectGDRFreeze a reference’s GDR space and project new data into it — the reference’s coordinates stay fixed.
Leiden at scale200,000 cells clustered in three seconds, with the same labels on any number of threads.
LeidenLocalRe-clustering inside one cluster, without re-running the whole analysis.
GDR at scale200,000 cells in 17 minutes, streamed — what the embedding costs at that size.
GDR and SVD on 1.5 million cellsBoth embeddings over one human cortex atlas: runtime, peak memory, and whether either separates cell types better.
GDR on developmental dataGroups that are stages rather than terminal cell types: 3× better separation for 1.6% accuracy.
GDR beyond transcriptomicsImages as expression matrices — what GDR separates on data that is not single-cell at all.
EmergeneCondition-specific gene programs scored per cell rather than per cluster, so heterogeneous responses stay visible.

Marker genes

what it covers
COSG: marker genes and significanceCosine-specificity markers, the p-value columns added in v1.2.0, IQR normalisation for comparing across cell types, and the one way of using the p-values that is wrong.
COSG on a cytomeThe same markers streamed from disk, and the four output shapes the streaming path returns.
COSG across batchesbatch_key scores each batch separately and averages — 86% of markers hold, and the 14% that move name your most dissociation-sensitive cell types.
COSG on the GPUOne argument, measured across five matrix sizes: 2× in the useful range, slower below 10,000 cells.
COSG on spatial dataOrgan markers on a whole embryo section, plotted back into tissue space — the check dissociated data cannot give you.

Annotation

what it covers
Marker-based cell type predictionLabel cells from a reference’s markers, checked against 8,738 held-out cells across 20 types (95.4%) and against PIASOmarkerDB.
PIASOmarkerDB API clientQuery 36 curated marker studies from Python, and go from a gene list back to the cell types that match.

Gene sets and sequence

what it covers
Gene set scoring (PIASOscore)Scoring a gene set — or a whole pathway database — against matched control sets, with per-cell p-values.
KEGG and drug-target gene sets320 pathways and 659 drug target sets scored per cell, then asked which cell type each belongs to.
Motif analysisGenome → promoters → PWM scan → enrichment, and the background choice that decides the answer.

Spatial transcriptomics

what it covers
Xenium with tissue-image overlayClusters drawn over the morphology image, and selecting a region of interest.
Xenium into a cytome directlyStraight from the platform output to a cytome, with no AnnData in between.
Xenium Prime 5K: mouse brain end to end63,173 cells × 5,006 genes from the raw bundle to annotated clusters, checked against known anatomy — and how to pull 83 MB out of a 13 GB archive.
Atera WTA18,028 targets in situ — 90% of the protein-coding transcriptome — and what that finds which a 313-gene panel cannot.
Downstream in situKEGG pathways per cell, LARIS ligand–receptor and cytorete regulons on both sections — and the coverage measurement that decides which are possible on your panel.
MERFISH sectionsSeveral sections in one cytome, aligned and analysed together.
Stereo-seq whole embryoA whole-embryo section at bin resolution, including rotating the coordinates.
GDR on spatial transcriptomicsEight embryonic stages, 520,815 bins, in one embedding.

Gene regulatory networks

what it covers
cytorete: RNA regulon inferenceTF → target regulons inferred from expression, constrained by promoter motifs.
cytorete on an AnnDataThe same chain in memory rather than on a cytome, and where each result lands.
cytorete on spatial dataRegulons across a whole embryo section, mapped back onto tissue.
Regulon dynamics across developmentHalf a million bins, eight stages, and which regulons move.

Cell-cell interaction

what it covers
SCALARLigand-receptor pairs between cell types, with a KNN-matched permutation null. 1.7M interactions in under a minute.
LARISThe spatial counterpart: when cells have coordinates, proximity constrains which interactions are possible. Links the six worked tutorials that ship with the package.

Plotting and data

what it covers
Plottingpiaso.pl end to end: embeddings, dot plots, violins, splits.
Colour palettesThe built-in palettes, and how to set your own.
Datasets and genome referencespiaso.data: what is available, and how caching works.
cytome basicsThe file format itself — what it stores and how to read it.
Converting: AnnData, Seurat, SingleCellExperimentcytome as an interchange format, and what does and does not travel with it.
cytome in RRead, write and stream .cytome natively in R, with no Python runtime.
Agents and project toolingPIASO-for-agents, stato and PlanDrop — running long analyses with coding agents.

Which one first?

Start with Human PBMC scRNA-seq end to end (AnnData) if your data is human, and Mouse brain scRNA-seq end to end (AnnData) if it is mouse. Either introduces every function the others reuse; the pair differs in what the QC step has to do, which is the part that does not transfer between samples. Median mitochondrial content is 12.0% in the human PBMC sample and 0.011% in the mouse nuclei, so on the human page every threshold bites and one cluster turns out to be mitochondrial reads, while on the mouse page the same thresholds are inert and the highest-mitochondrial cluster is endothelium that should be kept.

Then:

  • more cells than fit in memory → the cytome version of the same page;
  • more than one library → multiple samples;
  • clusters in hand and markers wanted → COSG.

Each AnnData page has a cytome twin running deliberately the same analysis. Reading a pair side by side is the fastest way to see what changes when the matrix stays on disk: almost nothing in the calls, everything in where the results live.

From the previous release

Thirteen more tutorials are published under /tutorials/previous/. They are listed here rather than only counted, because several cover ground no current page does — CellRanger pre-processing, and GDR on scATAC-seq.

what it covers
PIASO overviewThe v1.1.0 tour of the package, end to end.
INFOG + GDR on one million cellsThe Asian Immune Diversity Atlas at a million cells, before the streaming backend existed.
Processing: PBMC SAN2One PBMC sample from counts to annotated clusters.
Processing: PBMC SAN1 + SAN2The same, with two samples and the batch question that raises.
GDR on scATAC-seqGDR over chromatin rather than expression.
SCALAR: ligand-receptor analysisThe v1.1.0 SCALAR walkthrough on SEA-AD.
Marker-based cell type predictionThe earlier prediction workflow on mouse cortex.
PIASOmarkerDB API clientQuerying the marker database from Python, v1.1.0 style.
KEGG and ChEMBL gene setsPathway and drug-target sets, with ChEMBL alongside KEGG.
CellRanger: mouse brainFASTQ to count matrix for a mouse brain run.
CellRanger: mouseThe general mouse pre-processing recipe.
CellRanger: E18Pre-processing the E18 brain dataset used elsewhere in these pages.
CellRanger: E18 mouseThe E18 run against the mouse reference.

They were recovered from the PIASO v1.1.0 documentation and have not been re-run against this release, so each carries a banner saying so. They are worth reading for the analysis; check any call against the API reference before relying on it.

The original v1.1.0 site remains available at https://genecell.github.io/PIASO/ if you need the pages exactly as they were published, with their original outputs.