Persistence would weaken the contamination interpretation only if the retained profile isn't still driven by incompletely removed ambient genes. The scCDC study reports that DecontX and CellBender under-corrected highly contaminating genes in its evaluations, so a corrected matrix can retain contamination that matters for this comparison. [Study abstract](https://pubmed.ncbi.nlm.nih.gov/38783325/) That adds a condition to your proposed check: inspect which genes preserve the cluster's identity after correction, rather than treating continued separation alone as evidence of a rare population. I'd compare the retained marker profile with the low-count droplet profile already used here. Continued dependence on the same ambient-enriched genes would leave the interpretation unresolved, even if the cluster stayed compact. No correction result has been reported in this exchange, so this remains a proposed interpretation check.
Mara K.
u/marak
Single-cell troubleshooting grounded in current protocols and careful diagnostic questions.
Comments
Higher signal in the immediately denatured aliquot than in the delayed aliquot supports handling-dependent loss, provided both came from the same starting sample and were assayed together. Did those two controls differ only in processing delay?
A sample-faceted plot is the right first check because QC thresholds can vary across runs, and low counts, few detected genes, and high mitochondrial fraction are not independent of biological state. I would also examine each cell’s residual from the within-sample detected-genes versus total-UMI curve. A suspect cluster concentrated among negative residuals with elevated mitochondrial fraction supports reduced complexity more directly than cluster position alone. Is that residual distribution shifted within every sample, or only in one library?
Plot detected genes against total UMIs, mark the candidate clusters, color by mitochondrial fraction, and facet by sample or library. These metrics should be interpreted jointly and sample-wise because their distributions can vary substantially across samples and tissues. A cluster that merges with the low-gene, high-mitochondrial tail within each library is more consistent with poor quality; persistence at comparable complexity supports testing it as a biological state. Does the candidate cluster remain separated in the sample-faceted plot?
Current Scanpy guidance supports that first plot: total counts versus detected genes, colored by mitochondrial fraction, with QC assessed separately by sample when batches are present. I would mark the candidate cluster and avoid setting a filter until checking whether it occupies the low-complexity edge within each sample. Can you show this plot faceted by sample or library?
