Back to search

Article

Benchmarking Large Language Model Rationality Using Measurement Axioms

2026-07-01

Abstract excerpt

<title>Abstract</title> <p>While Large language models (LLMs) are increasingly deployed as decision-makers, current evaluation practice emphasizes benchmark accuracy rather than the structural properties that make decisions coherent and interpretable. We introduce a measurement-theoretic framework for evaluating AI decision-making using formal axioms derived from utility theory, focusing on transitivity of prefer...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
952df75c-fc87-5564-a8e4-9af4103f79f9
DOI
10.21203/rs.3.rs-10119008/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking Large Language Model Rationality Using Measurement AxiomsDOI 10.21203/rs.3.rs-10119008/v1
Select a neighboring publication to make it the new centre.