Back to search

Article

Mind the Gap: From Plausible to Valid Self-Explanations in Large Language Models

2025-04-02

Abstract excerpt

<title>Abstract</title> <p>This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations (SE) – extractive and counterfactual – using state-of-the-art LLMs (1B to 70B parameters) on two different classification tasks (objective and subjective). In line with Agarwal et al. (2024), our...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
c8d373c8-962e-57cd-8fb9-f360960cfe65
DOI
10.21203/rs.3.rs-6263278/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Mind the Gap: From Plausible to Valid Self-Explanations in Large Language ModelsDOI 10.21203/rs.3.rs-6263278/v1
Select a neighboring publication to make it the new centre.