Article
Graduated Dissent: Budgeted Disagreement Resolution for Multi-Model Inference
2026-03-24
Abstract excerpt
Recent empirical work demonstrates that large language models cannot reliably self-correct reasoning without external feedback, and large-scale evaluation across hundreds of models reveals substantial error correlation even between models with distinct architectures and providers. When generator and evaluator share failure modes, self-evaluation may provide weak evidence of correctness, and repeated self-critique...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- a517f4f3-c1ab-5f17-8f69-aca8cb002db6
- DOI
- 10.20944/preprints202603.1830.v1
