Article
Benchmarking Large Language Model Rationality Using Measurement Axioms
2026-07-01
Abstract excerpt
<title>Abstract</title> <p>While Large language models (LLMs) are increasingly deployed as decision-makers, current evaluation practice emphasizes benchmark accuracy rather than the structural properties that make decisions coherent and interpretable. We introduce a measurement-theoretic framework for evaluating AI decision-making using formal axioms derived from utility theory, focusing on transitivity of prefer...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 952df75c-fc87-5564-a8e4-9af4103f79f9
- DOI
- 10.21203/rs.3.rs-10119008/v1
