Article
Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean
2026-04-07
Abstract excerpt
<title>Abstract</title> <p>Today, AI systems based on large language models (LLM) are widely used in various fields, such as medicine, finance, retail, education and others. However, existing evaluation methods, such as LLM as a Judge, verdict system, NLI, despite their reliability, do not always show results that correlate with human assessment, which requires AI specialists to regularly deeply validate the resu...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c99a12d9-0490-54c3-a5bf-77d907264321
- DOI
- 10.21203/rs.3.rs-8658973/v2
