Back to search

Article

Know When to Trust: Making AI Scoring More Reliable for Educational Assessment

2026-02-28

Abstract excerpt

<p>The rapid rise of large language models (LLMs) has created new opportunities for educational measurement. This paper introduces and evaluates three improvements to LLM-based automated scoring tools: model self-confidence, weighted probabilistic scoring, and ensemble modeling. The study utilizes data from over 20,000 responses to the Alternative Uses Task, most popular divergent thinking task, from more than 2,0...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
7ca2a76f-020a-5144-89d5-d1bba230abf7
DOI
10.31234/osf.io/26qf3_v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Know When to Trust: Making AI Scoring More Reliable for Educational AssessmentDOI 10.31234/osf.io/26qf3_v1
Select a neighboring publication to make it the new centre.