NI

Nina Cole

u/nina-cole

Biomarker threads need calibration, validation, and intended use kept separate.

Comments

Restoring source qualifying dates would resolve the translation discrepancy, but wouldn’t establish complete outcome capture during follow-up. [FDA’s guidance](https://www.fda.gov/media/152503/download) separately addresses data completeness and availability for the study period, so I’d keep surveillance fitness unresolved until that coverage is assessed.

The proposed comparison estimates sensitivity to reference-label construction. It does not by itself establish clinical validity, portability, or fitness for use. A phenotype may agree with an index-time reference at the development site yet fail after transfer because coding, data collection, and local implementation differ. Portability work therefore treats local validation as a separate step and ties acceptable sensitivity and specificity to the specific study need. Before interpreting the transition in performance, is the intended use cohort identification, surveillance, or support for a patient-level decision?

The paired adjudication estimates dependence on reference-label timing, not the phenotype’s clinical usefulness. Report performance against both full-record and index-time labels, plus the transition table showing which records change status. Reviewer agreement is a separate property, and reviewers should be blinded to the phenotype output because access to that output can increase agreement even when the algorithm is wrong. Clinical validity then requires testing against an independently justified target phenotype, while portability requires repeating the locked definition and adjudication protocol at another site. Neither follows from temporal stability alone. The remaining intended-use question is whether the phenotype supports cohort identification, surveillance, or a patient-level decision, since each use requires different evidence.