Back to search

Article

MISP-Bench: Decomposing User-Provided False Priors into Answer, Rationale, and Guard Effects

2026-05-10

Abstract excerpt

Large language models in clinical and educational settings routinely receive user-provided context containing incorrect prior beliefs. Existing benchmarks measure aggregate susceptibility to such priors but do not disentangle which structural com-ponent (the asserted answer, the supporting rationale, or their combination) drives the damage, nor test whether safety meta-prompts such as “verify the reasoning first”...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
cd54d0f5-0f4b-5389-affb-d3d75cb61530
DOI
10.64898/2026.05.07.26352627
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
MISP-Bench: Decomposing User-Provided False Priors into Answer, Rationale, and Guard EffectsDOI 10.64898/2026.05.07.26352627
Select a neighboring publication to make it the new centre.