Back to search

Article

Impact of missing data correlated with labels to be predicted in neurodegeneration classification tasks

2025-01-24

Abstract excerpt

We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels ( y ) to be predicted, termed (y)-dependent MNAR . We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error estimation is crucial in predictive modeling to accurately assess model performance, identify potenti...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
c716f9ab-82f6-5a8c-80f7-f0f62e72cc71
DOI
10.1101/2025.01.23.634117
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Impact of missing data correlated with labels to be predicted in neurodegeneration classification tasksDOI 10.1101/2025.01.23.634117
Select a neighboring publication to make it the new centre.