Back to search

Article

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

2026-03-20

Abstract excerpt

<title>Abstract</title> <p>Today, AI systems based on large language models (LLM) are widely used in various fields, such as medicine, finance, retail, education and others. However, existing evaluation methods, such as LLM as a Judge, verdict system, NLI, despite their reliability, do not always show results that correlate with human assessment, which requires AI specialists to regularly deeply validate the resu...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
0c6679f0-dcd9-5aaf-98da-09bb32af145c
DOI
10.21203/rs.3.rs-8658973/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power MeanDOI 10.21203/rs.3.rs-8658973/v1
Select a neighboring publication to make it the new centre.