Baseline STARsolo mapping metrics (ce11/WS298) — 13 worm datasets, pre-3′UTR-extension #78
lhqing
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
STARsolo mapping metrics for the 13 C. elegans single-cell datasets (14 pipelines, 138 samples), mapped against the stock ce11 / WS298 annotation — the before snapshot. Once we build an extended 3′-UTR GTF and re-quantify, we compare against these numbers to answer #52: does the gene count go up?
What to watch
A 3′-UTR extension changes feature assignment, not alignment, so the informative columns are:
total_genes— genes detected per sample → expect ↑ if extended UTRs recover 3′ readsmed_gene— median genes per cell → expect ↑feat_uniq— fraction of reads uniquely assigned to a gene feature → the direct lever (reads landing just past the current 3′ end get captured)genome_UM— genome mapping → expected flat (same alignment; only the GTF's feature intervals change). This is the control: any gene-count gain must come from annotation, not re-mapping.Because alignment is unchanged, the re-quantification can likely run from the saved CRAMs without re-aligning (pending CB/UB tag retention — see #52).
Run config
genome ce11, annotation WS298, STARsolo;
soloFeatures = Gene, GeneFull, GeneFull_ExonOverIntron, GeneFull_Ex50pAS, Velocyto. Primary feature = Gene (scRNA) / GeneFull (snRNA). All 14 pipelines finishedEXIT=0; every sample produced.h5ad+.velocyto.h5ad+.cram+.qc.json.gz. Metrics are read from each sample'sqc.json.gz; column legend + flag thresholds are at the bottom.Baseline health: 0 SEVERE issues across 138 samples — 2 WARN (thin/low-input libraries) + 5 NOTE (shallow or undersequenced), all data-quality, none a pipeline fault. Machine-readable
mapping_metrics.tsvaccompanies this table.Flagged samples (ranked)
GSE126954 — CLEAN
10x v2 (sci-RNA-seq origin) · Gene — 7/7 samples
sci-RNA-seq origin mapped as 10x-v2: valid barcodes 0.94–0.98 across all 7 samples confirm the barcodes hit the v2 whitelist cleanly, so the chemistry resolution holds. SAMN10988847 called only 507 cells at ~228k reads/cell — a cell-calling knee that landed high (deep, few cells), worth an eye downstream but barcodes/mapping are fine. (8 GEO GSMs → 7 samples: the over-length experiment was absorbed into v2, expected.)
GSE136049 — MINOR NOTES
10x v2 + v3 (2 assays) · Gene · FACS-sorted neurons — 17/17 samples
CeNGEN FACS-sorted single-neuron-class libraries: the high saturation (~0.9–0.99) and large unique-vs-multi mapping gaps are the expected low-diversity, deeply-sequenced profile, not faults. INDEPENDENTLY VERIFIED (raw barcodes/features stats read on the node before it dropped).
GSE208154 — CLEAN
10x v3 (from possorted BAM) · Gene — 11/11 samples
Reconstructed from 10x possorted BAM (bamtofastq), multi-flowcell/multi-lane; was blocked as issue #30 and unblocked by fix F1. All 11 samples present with valid barcodes 0.96–0.98 — F1 seated the barcode read correctly and the index (I1) reads were dropped as intended.
GSE208229 — CLEAN
10x v3 · GeneFull · snRNA — 31/31 samples
snRNA (GeneFull): low fraction-of-reads-in-cells (0.10–0.44) is the expected high-ambient single-nucleus signature, not a defect. Many samples are moderate-nuclei / very-deeply-sequenced (sorted nuclei). Sample count verified 31/31 (units.tsv + results + rows) before the node dropped.
GSE226715 — MINOR NOTES
10x v3 · Gene · scRNA (shallow) — 3/3 samples
Whole-body N2 scRNA, deliberately shallow (30–67M reads): saturation 0.09–0.12 just means undersequenced — genome mapping (0.72–0.79) and per-cell complexity are otherwise healthy. A depth note, not a mapping fault.
GSE229022 — MINOR NOTES
10x v3 · GeneFull · snRNA (40 runs→28 samples) — 28/28 samples
snRNA aging atlas; 40 SRA runs correctly merge into 28 samples (verified units.tsv + results + rows before the node dropped). One low-input library (SAMN39550232, 221 cells) — barcodes 0.93 and genome mapping 0.95 are fine, so it's a thin library, not a pipeline fault.
GSE234962 — CLEAN
10x v3 · Gene · scRNA — 4/4 samples
Motor-neuron scRNA, all 4 samples healthy; a couple are oversequenced (saturation ~0.97) with deep cells — expected.
GSE256266 — CLEAN
10x v3 · GeneFull · snRNA (glia) — 4/4 samples
Glia snRNA (GeneFull): 2 of 4 samples (SAMN40022216/217) run a bit lower on genome mapping (0.73–0.78) and Q30 — read-quality-driven, still within usable range; barcodes 0.95.
GSE274290 — CLEAN
BD Rhapsody WTA · Gene — 1/1 samples
The sole BD Rhapsody WTA dataset: valid barcodes 0.897 (normal for BD's CLS whitelist), 14.9k cells, median 2933 genes/cell — the BD Rhapsody chemistry mapped correctly.
GSE310667 — MINOR NOTES
10x v3 · Gene · scRNA — 16/16 samples
Juvenile-neuron scRNA, all 16 present; a couple of shallow/low-complexity libraries (SAMN53327904/905, median 138–283 genes/cell) — thin cells, not mapping faults. Multimapping gaps are the low-complexity neuron profile.
PRJNA1027859 — CLEAN
10x v3 · GeneFull · snRNA (demo) — 6/6 samples
snRNA neuron atlas (seqforge demo), all 6 healthy; GeneFull-primary as intended.
PRJNA1195922 — CLEAN
10x v3 · GeneFull · snRNA (male) — 8/8 samples
Male-neuron snRNA, all 8 present; larger unique-vs-multi gaps and lower feat_uniq are the low-complexity/repeat neuron profile, within snRNA range. (SAMN49549038 sex=male is a metadata fix, F4 — not a QC signal.)
PRJNA658829 — CLEAN
10x v3 · Gene · scRNA (eQTL; 6 runs→2 BioSamples) — 2/2 samples
CRITICAL CHECK — RESOLVED. Both samples show valid barcodes 0.96/0.97, so fix F1 seated the barcode read on
_2correctly; the earlier near-zero-barcode concern does NOT materialize. 6 runs group into 2 BioSamples (SAMN15970313 ~2.2B reads; SAMN15970314 lower median genes/cell — a thinner library, but barcodes are valid).Appendix — all 138 samples (concatenated)
Full machine-readable table:
mapping_metrics.tsv(same directory). Columns: dataset, assay, sample, feat(ure), n_reads, valid_bc, genome_UM/uniq, feat_uniq, saturation, q30_cbumi/rna, est_cells, mean_reads_cell, median_umi/gene_cell, total_genes, frac_reads_cells.Flag thresholds: SEVERE = valid_bc<0.40 or genome_UM<0.30; WARN = valid_bc<0.70, genome_UM<0.60, feat_uniq<0.20, or cells<300; NOTE = median_gene<300 (shallow) or saturation<0.12 (undersequenced). snRNA (GeneFull) datasets aren't flagged for low feat_uniq or high ambient, which are expected for single-nucleus.
All reactions