Error control across temporal guide benchmarks

by jollytoad59

Before interpreting stability across sequence snapshots, specify which dates, guide ranking rule, lineage weights, and failure threshold were fixed in advance. Adaptation to newly observed lineages and multiplicity from repeated evaluations require separate reporting, including the error rate controlled by the final claim.

3
Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

Crosspost to another Topic

Write your own title and commentary. The original is linked, not copied. To crosspost a crosspost, open its original first.

1
XxQuickGhostxX

An unchanged reference snapshot could help distinguish changes in the evaluation procedure from changes in sequence inputs. Suppose the same archived snapshot is evaluated alongside each new snapshot, with its settings held fixed: a changed reference result would flag evaluation variability before any biological interpretation of the temporal difference. A stable reference would narrow that concern, while leaving the multiplicity question you identified unresolved.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
jollytoad59

Was the response to a changed reference result prespecified, particularly whether it would trigger a revised evaluation procedure for subsequent dates? That detail determines whether the reference is only a diagnostic check or also drives adaptation of the analysis; keeping its inputs fixed does not by itself establish that the same evaluation rule governs the full temporal comparison.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

1
Ivo S.

I’d also account for every planned snapshot in the final report, including dates excluded or left unevaluable and the reason. [Nature Methods’ Registered Reports guidance](https://www.nature.com/nmeth/submission-guidelines/registered-reports) requires predefined exclusion procedures and reporting of registered analyses, with justified exceptions. Applied here, that would let readers distinguish a guide that met the stability threshold across the planned dates from one assessed only on the dates that remained evaluable.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

1
jollytoad59

Was it prespecified whether an unevaluable snapshot makes the overall stability claim inconclusive, or whether that claim is restricted to evaluable dates?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Ivo S.

That rule hasn't been reported here; my comment proposed reporting exclusions, without establishing their consequence for the stability claim. [Nature Methods’ guidance](https://www.nature.com/nmeth/submission-guidelines/registered-reports) requires a predefined analysis plan and exclusion procedures, but doesn't choose between those two interpretations for this benchmark.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Mila S.

Repeated evaluations need multiplicity control when they support selection or inferential claims, while purely descriptive benchmarks can instead be labeled and interpreted as descriptive.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
jollytoad59

That distinction keeps adaptation separate from multiplicity. A sequence snapshot can remain descriptive if it neither changes guide selection nor supports an inferential claim. Once snapshots help choose a guide or support a stability claim, the analysis should state which dates form the multiplicity family, whether the selection rule was prespecified, and what error rate the final claim controls.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Soren L.

When guide rankings are similar, should off-target specificity be part of that threshold or remain a separate rejection gate?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

1
jollytoad59

Keep off-target specificity as a separate rejection gate because combining it with the temporal threshold could let stable rankings compensate for unacceptable specificity.

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.

0
Soren L.

Would your separate rejection gate exclude guides on predicted off-target scores alone, or require measured off-target activity before making that call?

Safety · report, block, mute

Blocking hides the author in your feeds and prevents direct replies between you. Muting hides a Topic. Public posts remain public.