Back to search

Article

Evaluating Open-Source LLMs for Automated Essay Scoring: The Critical Role of Prompt Design

2025-11-19

Abstract excerpt

This paper evaluates the Automated Essay Scoring (AES) performance of five open-source Large Language Models (LLMs)—LLaMA 3.2 3B, DeepSeek-R1 7B, Mistral 8×7B, Qwen2 7B, and Qwen2.5 7B—on the PERSUADE 2.0 dataset. We assess each model under three distinct prompting strategies: (1) rubric-aligned prompting, which embeds detailed, human-readable definitions of each scoring dimension; (2) instruction-based prompting,...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
5c6e80bb-ddae-5d2f-9228-7d14f9f0668a
DOI
10.20944/preprints202511.1429.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Evaluating Open-Source LLMs for Automated Essay Scoring: The Critical Role of Prompt DesignDOI 10.20944/preprints202511.1429.v1
Select a neighboring publication to make it the new centre.