Article
MISP-Bench: Decomposing User-Provided False Priors into Answer, Rationale, and Guard Effects
2026-05-10
Abstract excerpt
Large language models in clinical and educational settings routinely receive user-provided context containing incorrect prior beliefs. Existing benchmarks measure aggregate susceptibility to such priors but do not disentangle which structural com-ponent (the asserted answer, the supporting rationale, or their combination) drives the damage, nor test whether safety meta-prompts such as “verify the reasoning first”...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- cd54d0f5-0f4b-5389-affb-d3d75cb61530
- DOI
- 10.64898/2026.05.07.26352627
