Article
Risks of Using Large Language Models in Grading: LLMs and Humans Prefer LLM-Generated Writing Over Human’s but LLMs Show a Stronger Systematic Bias
2026-08-19
Abstract excerpt
<p>Recent studies suggest that large language models (LLMs) can approximate human grading performance, raising the prospect of their use in educational assessment. However, an important question remains: can LLMs evaluate student work fairly? In Study 1, we examined whether LLM evaluators exhibit systematic preference when grading psychology dissertations from three sources: (a) dissertations authored by 1,426 und...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- dd6418d8-9c4f-5f31-801a-1ead0a4b87f0
- DOI
- 10.31234/osf.io/35utw_v1
