Article
Impact of missing data correlated with labels to be predicted in neurodegeneration classification tasks
2025-01-24
Abstract excerpt
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels ( y ) to be predicted, termed (y)-dependent MNAR . We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error estimation is crucial in predictive modeling to accurately assess model performance, identify potenti...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c716f9ab-82f6-5a8c-80f7-f0f62e72cc71
- DOI
- 10.1101/2025.01.23.634117
