Article
Stress-Testing the Reasoning Competence of LLMs with Proofs Under Minimal Formalism
2026-04-23
Abstract excerpt
We introduce PROOFGRID, a challenging benchmark suite for evaluating the reasoning competence of language models through machine-checkable proofs rather than final answers alone. PROOFGRID spans 15 tasks organized around proof writing, proof checking, proof masking, and proof gap-filling. The tasks are expressed in deliberately minimal formal notation, most notably NDL, a stripped-down language for natural deducti...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9513fe01-399f-5c6d-9c04-6a2028207323
- DOI
- 10.20944/preprints202604.1702.v1
