Back to search

Article

Are LLMs better expert judges? Rethinking content validity assessment in the age of AI

2026-01-02

Abstract excerpt

<p>In this article, we demonstrate a novel application of large language models (LLMs) as expert judges for item-level content relevance evaluation. Eleven advanced LLMs were included, each treated as a separate expert panel based on multiple procedurally independent judgments generated via repeated API queries. Their performance was compared with ratings provided by human judges, including psychology students and...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
53a3290e-e6bc-5453-9c7b-360028707a13
DOI
10.31234/osf.io/4hsyx_v2
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Are LLMs better expert judges? Rethinking content validity assessment in the age of AIDOI 10.31234/osf.io/4hsyx_v2
Select a neighboring publication to make it the new centre.