Article
Are LLMs better expert judges? Rethinking content validity assessment in the age of AI
2026-01-02
Abstract excerpt
<p>In this article, we demonstrate a novel application of large language models (LLMs) as expert judges for item-level content relevance evaluation. Eleven advanced LLMs were included, each treated as a separate expert panel based on multiple procedurally independent judgments generated via repeated API queries. Their performance was compared with ratings provided by human judges, including psychology students and...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 83c756bc-d1eb-5b1a-a467-ba04544033f2
- DOI
- 10.31234/osf.io/4hsyx_v1
