Back to search

Article

Explainable Automatic Evaluation: Interpretable Decision-Making in Large Language Model Autoraters

2026-05-05

Abstract excerpt

The deployment of Large Language Models (LLMs) as automatic evaluators—termed "LLM autoraters"—has emerged as a scalable alternative to human judgment for assessing natural language generation quality. However, the opacity of LLM decision-making processes fundamentally undermines trust, debuggability, and regulatory compliance in high-stakes evaluation contexts. This study introduces a framework for explainable au...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
d7dce200-0f38-5025-a42f-37034b358b76
DOI
10.14293/pr2199.003509.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Explainable Automatic Evaluation: Interpretable Decision-Making in Large Language Model AutoratersDOI 10.14293/pr2199.003509.v1
Select a neighboring publication to make it the new centre.