Article
Benchmarking Large Language Models and Clinicians Using Locally Generated Primary Healthcare Vignettes in Kenya
2025-10-27
Abstract excerpt
<h4>Background</h4> Large language models (LLMs) show promise on healthcare tasks, yet most evaluations emphasize multiple-choice accuracy rather than open-ended reasoning. Evidence from low-resource settings remains limited. <h4>Methods</h4> We benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- d5c38d83-032e-5b49-8dd8-86922ec47147
- DOI
- 10.1101/2025.10.25.25338798
