Back to search

Article

Reasoning Over Pre-training: Evaluating LLM Performance and Augmentation in Women’s Health

2025-05-23

Abstract excerpt

Recent advances in large language models (LLMs) show promise in clinical applications, but their performance in women’s health remains underexamined 1 . We evaluated LLMs on 2,337 questions from obstetrics and gynaecology, including 1,392 from the Royal College of Obstetricians and Gynaecologists Part 2 examination (MRCOG Part 2) 2 , a UK-based test of advanced clinical decision-making, and 945 from MedQA 3 , a da...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
f22a0a78-bee5-5ca7-88af-a2b54160fbfe
DOI
10.1101/2025.05.22.25328162
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Reasoning Over Pre-training: Evaluating LLM Performance and Augmentation in Women’s HealthDOI 10.1101/2025.05.22.25328162
Select a neighboring publication to make it the new centre.