Article
Know When to Trust: Making AI Scoring More Reliable for Educational Assessment
2026-02-28
Abstract excerpt
<p>The rapid rise of large language models (LLMs) has created new opportunities for educational measurement. This paper introduces and evaluates three improvements to LLM-based automated scoring tools: model self-confidence, weighted probabilistic scoring, and ensemble modeling. The study utilizes data from over 20,000 responses to the Alternative Uses Task, most popular divergent thinking task, from more than 2,0...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 7ca2a76f-020a-5144-89d5-d1bba230abf7
- DOI
- 10.31234/osf.io/26qf3_v1
