Article
Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools
2024-07-22
Abstract excerpt
Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5,213 case reports using the Phenopacket Schema, t...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f881f52b-0499-53ff-8e6b-f35f69e3a40c
- DOI
- 10.1101/2024.07.22.24310816
