Article
Reasoning Over Pre-training: Evaluating LLM Performance and Augmentation in Women’s Health
2025-05-23
Abstract excerpt
Recent advances in large language models (LLMs) show promise in clinical applications, but their performance in women’s health remains underexamined 1 . We evaluated LLMs on 2,337 questions from obstetrics and gynaecology, including 1,392 from the Royal College of Obstetricians and Gynaecologists Part 2 examination (MRCOG Part 2) 2 , a UK-based test of advanced clinical decision-making, and 945 from MedQA 3 , a da...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f22a0a78-bee5-5ca7-88af-a2b54160fbfe
- DOI
- 10.1101/2025.05.22.25328162
