Back to search

Article

CalibJudge: Calibrated LLM-as-a-Judge for Multilingual RAG with Uncertainty-Aware Scoring

2026-03-17

Abstract excerpt

Large Language Models (LLMs) serving as automatic evaluators (LLM-as-a-Judge) have become essential for assessing Retrieval-Augmented Generation (RAG) systems. However, in multilingual settings, these judges exhibit significant calibration drift across languages, producing scores that are neither comparable nor aligned with human judgments. We present CalibJudge, a post-hoc calibration framework that addresses thi...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
51791edd-5439-5e89-b620-ce33b982c26d
DOI
10.20944/preprints202603.1324.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
CalibJudge: Calibrated LLM-as-a-Judge for Multilingual RAG with Uncertainty-Aware ScoringDOI 10.20944/preprints202603.1324.v1
Select a neighboring publication to make it the new centre.