Back to search

Article

A statistical framework for evaluating the repeatability and reproducibility of large language models

2025-08-08

Abstract excerpt

<h4>Objective</h4> Systematic evaluation of variability in large language model (LLM)-generated outputs is critical for assessing their reliability. We developed a regulatory-informed statistical framework to quantify LLM repeatability and reproducibility. <h4>Materials and Methods</h4> Using definitions from U.S. Food and Drug Administration draft guidance for AI-enabled medical software, our framework quantifi...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
b71575f4-e7b7-5c5e-a73e-f7da5757e094
DOI
10.1101/2025.08.06.25333170
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A statistical framework for evaluating the repeatability and reproducibility of large language modelsDOI 10.1101/2025.08.06.25333170
Select a neighboring publication to make it the new centre.