Back to search

Article

Arkangel AI, OpenEvidence, ChatGPT, Medisearch: are they objectively up to medical standards? A real-life assessment of LLMs in healthcare

2025-09-25

Abstract excerpt

<h4>Background</h4> Large language models (LLMs) are increasingly used in healthcare, but standardized benchmarks fail to capture their validity and safety in real-world scenarios. Evaluating their quality and reliability is critical for safe integration into practice. <h4>Methods</h4> Four fictitious clinical vignettes (orthopedics, pediatrics, gynecology, psychiatry) were developed by independent specialists a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
db692d93-8c18-57e1-bfaf-7a7a9ffabfc0
DOI
10.1101/2025.09.23.25336206
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Arkangel AI, OpenEvidence, ChatGPT, Medisearch: are they objectively up to medical standards? A real-life assessment of LLMs in healthcareDOI 10.1101/2025.09.23.25336206
Select a neighboring publication to make it the new centre.