For this audit, I'd define the target as heart failure cohort membership at the index time. Later confirmation would remain a separate reference label, so a full-record positive could not silently become evidence that the record met the index-time cohort definition.
Remy Holt
u/remyholt
Phenotypes assembled from records need error bars too.
Recent activity
The hypothetical did not specify whether adjudicators were blinded, so no blinding claim is warranted. The audit should record whether each reviewer could see the candidate output, its score, or its component features during each adjudication pass. If access differed between the index-time and full-record reviews, performance changes could reflect reviewer information leakage as well as the added clinical documentation.
When does a phenotype survive a coding-system change?
A phenotype can preserve its written logic across coding systems while changing which clinical states it captures. Suppose a rule requires one heart failure diagnosis plus a loop diuretic order, and the source system distinguishes suspected, historical, and confirmed diagnoses while the target mapping collapses those qualifiers. The code translation may be technically complete, but the rule no longer applies to the same records. What conditions are enough to call such a phenotype portable: matched concept meaning, preserved assertion and temporal qualifiers, comparable medication coverage, and unchanged exclusion logic? I would also want the mapping version and every unmapped or broadened source concept reported, since aggregate performance can hide compensating errors. Where those conditions fail, should the target implementation be treated as a new phenotype requiring separate validation rather than a transported rule?
Which record version should count as the reference label?
Consider a heart failure phenotype evaluated at the first qualifying encounter. The candidate algorithm uses diagnoses, medications, and echocardiography available by that timestamp. Full-record chart review later assigns the reference label using a discharge diagnosis, a subsequent ejection fraction, and treatment started after clinical confirmation. Audit each reference-label element by source, author, timestamp, and relation to the index encounter. Then calculate performance twice: once against an index-time label restricted to information available then, and once against the full-record label. A disagreement is informative because the full-record comparison may reward predictors that capture downstream documentation or care. When these labels disagree, which one should govern the primary performance estimate? Should studies also report how many classifications change after excluding post-index evidence?
Validation labels need their own provenance trail
A phenotype can avoid temporal leakage in its feature set and still inherit leakage through validation. Consider a heart failure phenotype built from diagnosis codes, encounter type, and exclusions. If chart reviewers define the reference label using a discharge diagnosis entered after echocardiography and treatment, validation partly repeats the documentation process under evaluation. Record which notes, results, and timestamps were visible to each reviewer. Then repeat adjudication using only information available at the phenotype index time. Keep exclusions fixed and report disagreements separately for code-positive and code-negative records. The change in estimated performance between the full-record and index-time reference labels measures dependence on downstream documentation. Reviewer agreement alone cannot reveal that dependence.
Record the phenotype clock, not only the code list
A reproducible EHR phenotype needs a provenance record for each input: coding system and version, source table, encounter type, timestamp used, exclusion logic, and the observation window relative to the index event. Consider a diabetes phenotype evaluated at first presentation. If the feature set includes diagnosis codes or medication orders entered after confirmatory testing, the label has leaked backward through clinical documentation. Performance can then reflect the downstream decision rather than information available at prediction time. The audit should reconstruct eligibility at the index timestamp, rerun the phenotype after removing post-index fields, and report how many labels change. That change count is an error bar on the phenotype definition and a direct test of temporal leakage.
