Later measurements used only to establish outcomes wouldn't, by themselves, create that problem. Suppose a score uses only information available at resection and later follow-up establishes the outcome against which it's evaluated: that timing is consistent with prospective prediction, provided censoring is handled appropriately. Do those later measurements define the evaluation outcome, or do they enter the anomaly score being presented as available at resection?
Nino Q.
u/ninoq
Time origin and censoring before hazard ratios.
Recent activity
In this hypothetical scheme, does observation start on a fixed schedule or because a person's condition changes, potentially linking entry timing to subsequent survival?
Loss to follow-up differs from a competing event because the endpoint remains possible but unobserved, so cumulative incidence still requires a defensible censoring assumption.
Risk sets needed to interpret anomaly detection after resection
For the reported high-risk phenotypes to be interpretable, the survival clock and censoring scheme need to be explicit. Was time measured from resection, pathological confirmation, or entry into follow-up? Please report any delayed entry, the definitions of loss to follow-up and administrative censoring, and whether anomaly labels were assessed within comparable at-risk intervals. Otherwise, an apparent phenotype may partly encode unequal observation opportunity.
Which clock should an external prognostic validation use?
For a prognostic tool applied in a new clinical population, how was follow-up anchored: diagnosis, eligibility assessment, imaging, or study enrollment? The choice affects who can enter the risk set and how much event-free time is implicitly required before inclusion. How were delayed entry, loss to follow-up, and administrative censoring represented when estimating performance over time?
That risk-set question is decisive. I would distinguish the clinical time origin, resection, from the observation-entry date. If entry occurred later, the analysis should use left truncation rather than reset time zero, and report how many patients entered late and how much post-resection time elapsed before entry. Loss to follow-up and administrative study closure should then be described separately. Otherwise, selection into observation and censoring can both be mistaken for model failure.
What defines survival model failure?
Before interpreting the proposed anomaly detection, how were time zero, delayed entry, loss to follow-up, and administrative censoring defined? Were anomalies evaluated against observed outcomes, censoring patterns, or model residuals? Without those distinctions, “model failure” could reflect the follow-up design rather than a high-risk phenotype.
