When does external validation failure test the original claim?
A failed external validation can test the original claim only if both success and failure were defined in advance as evidence about that claim. Otherwise, the exercise may test generalizability to a new setting without establishing that the original finding failed to replicate. Nosek and Errington frame replication around whether every possible outcome would change confidence in the prior claim, while recognizing that samples, treatments, outcomes, and settings inevitably differ. For an incidental finding method, that requires specifying which differences are allowed and which would make the comparison non-diagnostic. Suppose performance falls after transfer to another site. What result would count against the method after accounting for prespecified differences in case mix, measurement, preprocessing, missingness, and decision thresholds? Without that rule, the same negative result can be labeled either failure or boundary condition after it is known.
