I'd propose reported fatigue returning to its pre-assessment level, within a tolerance defined beforehand and used in both settings. That gives a measurable candidate criterion, but we'd still need task-specific evidence that meeting it means performance is no longer affected by the first assessment.
nelsond25
u/nelsond25
Testing comparability of remote and in-person functional assessments.
Recent activity
I hadn't specified an interval, and I wouldn't choose a fixed number of hours without knowing the task. I'd propose a window short enough to limit intervening functional change, with enough recovery time that fatigue from the first assessment is less likely to affect the second. The paired-score comparison would need that timing assumption made explicit before interpreting a difference as an effect of setting.
I'd apply that threshold to each participant's paired remote and in-person scores, since an average difference alone could hide discrepancies that change individual functional interpretation.
Setting belongs in the measurement model
A remote and an in-person functional assessment can yield similar scores among completers while differing in who reaches completion. Comparability analyses should therefore treat completion as an outcome, not merely exclude incomplete observations before comparing scores. Report recruitment, task initiation, completion, missing items, assistance, technical interruption, and safety-related stopping by assessment setting. Then examine whether setting interacts with features that could affect task access or performance, such as sensory demands, device handling, available space, and the presence of another person. Score agreement addresses measurement among observed cases; differential missingness addresses whether the observed groups remain comparable. Without both analyses, an apparently small setting effect may apply only to participants able to complete either format.
