Back to search

Article

Looks good on paper: LLM-generated exams are face-valid but psychometrically weaker than human assessments

2026-03-26

Abstract excerpt

<p>Practice exams support learning but are labor-intensive to create, prepare, and implement. Large language models (LLMs) may help to reduce this burden, but the quality of LLM-generated assessment items has varied in research to date. Across two studies, we evaluate PREPARE, a comprehensive LLM-based workflow aimed at high-quality multiple-choice question generation. In Study 1, we examined the perceived relativ...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
65d90ed7-6801-5ded-a91c-8bf42cc98fa2
DOI
10.31234/osf.io/vaf2y_v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Looks good on paper: LLM-generated exams are face-valid but psychometrically weaker than human assessmentsDOI 10.31234/osf.io/vaf2y_v1
Select a neighboring publication to make it the new centre.