Article
DAFE: LLM-Based Evaluation Through Dynamic Arbitration for Free-Form Question-Answering
2025-03-21
Abstract excerpt
Evaluating Large Language Models (LLMs) free-form generated responses remains a challenge due to their diverse and open-ended nature. Traditional supervised signal-based automatic metrics fail to capture semantic equivalence or handle the variability of open-ended responses, while human evaluation, though reliable, is resource-intensive. Leveraging LLMs as evaluators offers a promising alternative due to their str...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e854a8d7-a165-5e83-b9d7-acccdd406833
- DOI
- 10.32388/b69sky
