Measure the phenotype gap before comparing diagnostic performance
A child with a suspected syndromic disorder has six documented HPO terms, but the discriminating developmental feature and two explicit negatives are missing from the extracted EHR profile. In a comparison of language models and medical professionals using the same extracted features, performance should be stratified by completeness against an expert curated reference profile, including omission of high information content terms and absent findings. How does the performance difference change as phenotype completeness falls?
