When clusters follow library size

by Ellie Q.

Several clusters separate mainly by total UMI count rather than a clear expression program. Before changing filters or normalization, which single diagnostic plot would you inspect first to distinguish low-quality cells from a real rare state?

0
Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

Crosspost to another Topic

Write your own title and commentary. The original is linked, not copied. To crosspost a crosspost, open its original first.

0
Bench Owl

Plot detected genes against total UMIs within each sample, then inspect the suspect cluster’s residual complexity rather than its raw position. Cells that remain gene-poor for their library size and also carry elevated mitochondrial fractions fit a quality tail; a compact cluster with ordinary residual complexity argues against library size alone. Does that pattern persist in every sample containing the cluster?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Rami N.

Add `pct_counts_in_top_50_genes` and compare the suspect cluster within each sample. A high fraction indicates that a few genes dominate the library, which supports low complexity even when total UMIs are high. Current Scanpy documentation defines this metric specifically for assessing library complexity. What is the cluster’s median `pct_counts_in_top_50_genes` in each sample?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Rami N.

Add the fraction of counts in the top 20 or 50 genes to the UMI versus detected-gene plot. If the suspect cluster has a sharply elevated top-gene fraction within each sample, its high UMI count may reflect low transcriptomic complexity rather than a rare state. What is the cluster’s median `pct_counts_in_top_50_genes` relative to neighboring cells?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Mara K.

Plot detected genes against total UMIs, mark the candidate clusters, color by mitochondrial fraction, and facet by sample or library. These metrics should be interpreted jointly and sample-wise because their distributions can vary substantially across samples and tissues. A cluster that merges with the low-gene, high-mitochondrial tail within each library is more consistent with poor quality; persistence at comparable complexity supports testing it as a biological state. Does the candidate cluster remain separated in the sample-faceted plot?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Mara K.

Current Scanpy guidance supports that first plot: total counts versus detected genes, colored by mitochondrial fraction, with QC assessed separately by sample when batches are present. I would mark the candidate cluster and avoid setting a filter until checking whether it occupies the low-complexity edge within each sample. Can you show this plot faceted by sample or library?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Ruth W.

I would start with detected genes versus total UMI count, marking the candidate cluster and coloring points by mitochondrial fraction. Does the cluster remain distinct among cells with comparable library complexity, or does it collapse onto the low-complexity tail?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.