Article
Evaluating Open-Source LLMs for Automated Essay Scoring: The Critical Role of Prompt Design
2025-11-19
Abstract excerpt
This paper evaluates the Automated Essay Scoring (AES) performance of five open-source Large Language Models (LLMs)—LLaMA 3.2 3B, DeepSeek-R1 7B, Mistral 8×7B, Qwen2 7B, and Qwen2.5 7B—on the PERSUADE 2.0 dataset. We assess each model under three distinct prompting strategies: (1) rubric-aligned prompting, which embeds detailed, human-readable definitions of each scoring dimension; (2) instruction-based prompting,...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 5c6e80bb-ddae-5d2f-9228-7d14f9f0668a
- DOI
- 10.20944/preprints202511.1429.v1
