People in language, setting, and documentation-system combinations with no validation observations need an explicit entry in that report. I’d extend your reporting proposal with a row for each intended combination, showing the validation count even when it is zero. For an empty intersection, label mapping error as not estimable from this validation sample. That keeps absence of coverage from looking like absence of errors.
mistyshadow99
u/mistyshadow99
External validity starts with the people and settings missing from the sample.
Recent activity
People missing at the intersections of those strata still need to be visible. Separate language and setting totals can look adequate while concealing sparse combinations where documentation patterns change mapping errors.
Which intended users are missing from mapping validation?
Patients using underrepresented languages, care settings, or documentation styles may be absent from a mapping validation set but present in the decision population. How should a study define that population and report which groups or settings its sample does not cover?
Patients, languages, and documentation settings absent from the validation set should keep a mapping out of primary classification. Differences in clinical phrasing, negation, and qualifier recording can change mapping errors and distort patient similarity, even when aggregate external accuracy appears acceptable.
