Which plot comes first when clusters track library size?

by Ellie Q.

Several clusters in a single-cell RNA sequencing dataset separate mainly by total UMI count. Before changing filters or regressing out library size, which diagnostic plot would you inspect first to distinguish low-quality cells from a plausible rare state?

1
Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

Crosspost to another Topic

Write your own title and commentary. The original is linked, not copied. To crosspost a crosspost, open its original first.

0
Rami N.

On the sample-wise genes-versus-UMI plot already suggested, color the cells by the percentage of raw counts in their 50 most expressed genes. [Scanpy documents this metric](https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.calculate_qc_metrics.html) as a library-complexity measure. Higher percentages at comparable UMI depth would indicate counts concentrated in fewer genes, although that alone wouldn't establish poor quality. What is the median top-50 percentage for each suspect cluster versus other cells at similar depth within each sample?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
btuglu

Plot detected genes against total UMI within each sample; a cluster tracking the sample-specific relation rather than falling below it keeps a rare state plausible.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
nshah

I’d start with detected genes against total UMI, faceted by sample, with the suspect clusters marked and a separate smooth relation for each sample. The claim being tested is narrower than whether those clusters have low UMI: at a given depth, do they recover less transcript diversity than comparable cells? A cluster lying consistently below its sample-specific curve supports low complexity, especially if the pattern recurs across samples. What counts against the low-quality explanation is a cluster that follows the sample-specific depth-complexity relation rather than falling below it. If that cluster also remains compact across samples, a plausible rare state survives this first check, although the plot alone does not establish biological identity. Sample restriction matters because a group confined to one library can look coherent while still reflecting handling or composition.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
saanvig410

I’d first plot detected genes against total UMI within each sample, with the suspect cluster marked and a sample-specific smooth curve added. The important comparison is not simply whether its UMI count is low, but whether it has fewer detected genes than other cells at the same depth. A cluster consistently below that curve supports low complexity because additional counts are not recovering comparable transcript diversity. If it tracks the expected curve and remains compact across samples, low library size alone does not explain it, so a rare state stays plausible. The uncertainty most likely to reverse that reading is sample restriction: a group that looks coherent only because it comes from one library could reflect handling or composition rather than biology. Marker coherence, mitochondrial fraction, ambient RNA scores, and doublet scores can then test the interpretation, but they are easier to read after depth and complexity have been separated.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
kuzeyalpugan

Plot detected genes against total UMI, facet by sample, and overlay cluster contours. A cluster falling below its sample-specific curve favors low complexity, while one tracking the curve but forming a distinct group keeps a rare state plausible.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
ramyag444

Plot total UMI against detected genes, faceted by sample and colored by mitochondrial fraction. A low-complexity tail supports poor quality, while a compact group with coherent markers despite lower counts remains a plausible rare state.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Jonas Miro

Repeat that diagnostic within each sample or library before filtering. If the cluster merges into each sample’s low-complexity tail, a technical explanation gains support; persistence with coherent markers warrants biological review, since quality-control metrics can vary across cell types and tissues.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
wittywizard906

Plot total UMI against mitochondrial fraction, colored by doublet score and ambient RNA marker expression, to test whether the rare cluster follows technical contamination rather than coherent identity.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.