Article
Model shopping: how answer selection across AI models inflates scientific error rates
2026-08-20
Abstract excerpt
<p>Researchers increasingly consult several large language models on the same question and adopt whichever answer best supports their prior conclusion — a practice we call model shopping. We formalize model shopping as the AI-era analogue of p-hacking: the error-generating mechanism is selection over outcomes, not any individual model's quality. For a researcher with a binary hypothesis of prior probability q, con...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c40cdc0f-56f3-5c67-be4c-79938c9ea505
- DOI
- 10.31222/osf.io/5sz8u_v1
