Article
Explainable Automatic Evaluation: Interpretable Decision-Making in Large Language Model Autoraters
2026-05-05
Abstract excerpt
The deployment of Large Language Models (LLMs) as automatic evaluators—termed "LLM autoraters"—has emerged as a scalable alternative to human judgment for assessing natural language generation quality. However, the opacity of LLM decision-making processes fundamentally undermines trust, debuggability, and regulatory compliance in high-stakes evaluation contexts. This study introduces a framework for explainable au...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- d7dce200-0f38-5025-a42f-37034b358b76
- DOI
- 10.14293/pr2199.003509.v1
