Article
Limits of Self-Correction in LLMs: An Information-Theoretic Analysis of Correlated Errors
2026-05-25
Abstract excerpt
We develop a diagnostic framework for evaluating when LLM self-evaluation can be trusted. The framework's central results are: (1) under a shared-blind-spot modeling assumption with joint conditional independence of evaluations given the shared failure structure, k rounds of self-critique provide information about correctness bounded by what the shared latent failure variable Z mediates---not by any independent ch...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- ae3b5b29-436f-502c-bba0-8c5d299566e4
- DOI
- 10.20944/preprints202601.0892.v5
