Back to search

Article

Can Large Language Models Detect Contradicted Biomedical Evidence? A Dual-Annotator Benchmark Audit

2026-07-06

Abstract excerpt

<title>Abstract</title> <p>Background: Biomedical question answering systems increasingly use retrieval-augmented generation, yet generated answers may be unsupported by, or actively contradict, retrieved evidence. In the configurations tested in this study, all LLM verifiers substantially under-detected gold contradictions, with the best automated recall only 5/23 (21.7%). We audited this failure mode and tested...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
e6bb362c-5bf7-5e12-9d17-8a4d629e1d5f
DOI
10.21203/rs.3.rs-10064275/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Can Large Language Models Detect Contradicted Biomedical Evidence? A Dual-Annotator Benchmark AuditDOI 10.21203/rs.3.rs-10064275/v1
Select a neighboring publication to make it the new centre.