Back to search

Article

A new pipeline for cross-validation fold-aware machine learning prediction of clinical outcomes addresses hidden data-leakage in omics based ‘predictors’

2026-03-16

Abstract excerpt

<h4>Motivation</h4> Machine learning (ML) approaches are increasingly applied to high-dimensional biological data in which features are often dataset-dependent. In many omics workflows, features are computed using information derived from the entire dataset, such as correlations between variables, clustering structures, or enrichment scores. We refer to these as global dataset features , defined as features whos...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
b3f7351a-ef00-54b2-98f2-6f3841877e62
DOI
10.64898/2026.03.12.711429
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A new pipeline for cross-validation fold-aware machine learning prediction of clinical outcomes addresses hidden data-leakage in omics based ‘predictors’DOI 10.64898/2026.03.12.711429
Select a neighboring publication to make it the new centre.