Back to search

Article

A Three-Tier Operational Benchmark for Evaluating Large Language Models on Hospital Medication Safety

2026-06-10

Abstract excerpt

<h4>Objective</h4> To introduce PsiBench, a clinically validated medication-safety benchmark for evaluating large language models (LLMs) against the standards used to certify hospital computerized provider order entry (CPOE) and electronic health record (EHR) systems, and a non-overlapping three-tier evaluation framework separating highest-stakes discrimination, the operational CDS regime, and category-correct al...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
ccf224b1-a2b0-545b-bf0c-4095a2a468a1
DOI
10.64898/2026.06.05.26354271
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A Three-Tier Operational Benchmark for Evaluating Large Language Models on Hospital Medication SafetyDOI 10.64898/2026.06.05.26354271
Select a neighboring publication to make it the new centre.