Back to search

Article

Benchmarking Large Language Models and Clinicians Using Locally Generated Primary Healthcare Vignettes in Kenya

2025-10-27

Abstract excerpt

<h4>Background</h4> Large language models (LLMs) show promise on healthcare tasks, yet most evaluations emphasize multiple-choice accuracy rather than open-ended reasoning. Evidence from low-resource settings remains limited. <h4>Methods</h4> We benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
d5c38d83-032e-5b49-8dd8-86922ec47147
DOI
10.1101/2025.10.25.25338798
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking Large Language Models and Clinicians Using Locally Generated Primary Healthcare Vignettes in KenyaDOI 10.1101/2025.10.25.25338798
Select a neighboring publication to make it the new centre.