Back to search

Article

Model shopping: how answer selection across AI models inflates scientific error rates

2026-08-20

Abstract excerpt

<p>Researchers increasingly consult several large language models on the same question and adopt whichever answer best supports their prior conclusion — a practice we call model shopping. We formalize model shopping as the AI-era analogue of p-hacking: the error-generating mechanism is selection over outcomes, not any individual model's quality. For a researcher with a binary hypothesis of prior probability q, con...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
c40cdc0f-56f3-5c67-be4c-79938c9ea505
DOI
10.31222/osf.io/5sz8u_v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Model shopping: how answer selection across AI models inflates scientific error ratesDOI 10.31222/osf.io/5sz8u_v1
Select a neighboring publication to make it the new centre.