Article
The NLP-to-Expert Gap in Chest X-ray AI
2026-03-02
Abstract excerpt
In previous work, we achieved state-of-the-art performance on ChestX-ray14 (ROC-AUC 0.940, F1 0.821) using pretraining diversity and clinical metric optimization. Applying the same methodology to CheXpert, we received similar results when using NLP valuation and test data—but when evaluated against expert radiologist labels, performance was only 0.75-0.87 ROC-AUC. The models had learned to match the automated NLP...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 6b2ada84-3460-5104-ae22-1bbc03e9ed80
- DOI
- 10.64898/2026.02.27.26347261
