Article
Mind the Gap: From Plausible to Valid Self-Explanations in Large Language Models
2025-04-02
Abstract excerpt
<title>Abstract</title> <p>This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations (SE) – extractive and counterfactual – using state-of-the-art LLMs (1B to 70B parameters) on two different classification tasks (objective and subjective). In line with Agarwal et al. (2024), our...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c8d373c8-962e-57cd-8fb9-f360960cfe65
- DOI
- 10.21203/rs.3.rs-6263278/v1
