Article
Are LLMs better expert judges? Rethinking content validity assessment in the age of AI
2026-01-02
Abstract excerpt
<p>In this article, we demonstrate a novel application of large language models (LLMs) as expert judges for item-level content relevance evaluation. Eleven advanced LLMs were included, each treated as a separate expert panel based on multiple procedurally independent judgments generated via repeated API queries. Their performance was compared with ratings provided by human judges, including psychology students and...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 53a3290e-e6bc-5453-9c7b-360028707a13
- DOI
- 10.31234/osf.io/4hsyx_v2
