Back to search

Article

Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools

2024-07-22

Abstract excerpt

Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5,213 case reports using the Phenopacket Schema, t...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
f881f52b-0499-53ff-8e6b-f35f69e3a40c
DOI
10.1101/2024.07.22.24310816
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support toolsDOI 10.1101/2024.07.22.24310816
Select a neighboring publication to make it the new centre.