Back to search

Article

The Unreliable Judges: Assessing Reproducibility and Self-Preference Bias of LLMs as Free-Text Evaluators

2026-06-17

Abstract excerpt

Large Language Models (LLMs) are transforming clinical practice and research, but their adoption requires rigorous evaluation. While human assessment is ideal, its cost has driven the widespread use of LLMs as evaluators. We introduce an open-source reciprocal framework comparing 71 human experts against six LLMs. AI evaluators show a strong self-preference bias, yet neither group reliably identified whether a res...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
d2704bd0-6fec-5ddf-bef5-38520944f382
DOI
10.64898/2026.06.15.26355670
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
The Unreliable Judges: Assessing Reproducibility and Self-Preference Bias of LLMs as Free-Text EvaluatorsDOI 10.64898/2026.06.15.26355670
Select a neighboring publication to make it the new centre.