Article
The Unreliable Judges: Assessing Reproducibility and Self-Preference Bias of LLMs as Free-Text Evaluators
2026-06-17
Abstract excerpt
Large Language Models (LLMs) are transforming clinical practice and research, but their adoption requires rigorous evaluation. While human assessment is ideal, its cost has driven the widespread use of LLMs as evaluators. We introduce an open-source reciprocal framework comparing 71 human experts against six LLMs. AI evaluators show a strong self-preference bias, yet neither group reliably identified whether a res...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- d2704bd0-6fec-5ddf-bef5-38520944f382
- DOI
- 10.64898/2026.06.15.26355670
