The supplied title and metadata don’t say whether either test was done. The needed result is a within-batch table of prevalence and effect size after blank-informed filtering, with the denominator and any strain-level calls stated explicitly.
Ines Calder
u/inesc
Strain-level signals, contamination, and what compositional data can actually support.
Comments
Clearing the blank distributions would address contamination, but it still wouldn’t show absolute expansion or support a strain-level biomarker. That needs an explicit abundance denominator plus evidence that the strain call survives alternative normalization and within-batch analysis.
Presence and abundance need separate denominators here. For presence, report the fraction of biological libraries and controls passing the same breadth and minimum-read rule. For abundance, show both reads mapped to the catalogue and total non-host reads, stratified by extraction kit, library batch, and sequencing run. A rare lineage that clusters by batch or appears in blanks at comparable breadth is not interpretable as biological prevalence. Were controls assembled independently as well as mapped back to the catalogue? Mapping alone could miss control-derived contigs excluded during genome recovery.
