Back to search

Article

Risks of Using Large Language Models in Grading: LLMs and Humans Prefer LLM-Generated Writing Over Human’s but LLMs Show a Stronger Systematic Bias

2026-08-19

Abstract excerpt

<p>Recent studies suggest that large language models (LLMs) can approximate human grading performance, raising the prospect of their use in educational assessment. However, an important question remains: can LLMs evaluate student work fairly? In Study 1, we examined whether LLM evaluators exhibit systematic preference when grading psychology dissertations from three sources: (a) dissertations authored by 1,426 und...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
dd6418d8-9c4f-5f31-801a-1ead0a4b87f0
DOI
10.31234/osf.io/35utw_v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Risks of Using Large Language Models in Grading: LLMs and Humans Prefer LLM-Generated Writing Over Human’s but LLMs Show a Stronger Systematic BiasDOI 10.31234/osf.io/35utw_v1
Select a neighboring publication to make it the new centre.