Article
Discovering Latent Sycophantic Patterns in LLMs via Distributional Anomaly Detection
2026-05-08
Abstract excerpt
Large language models (LLMs) often exhibit sycophancy—the tendency to generate responses that align with a user’s stated beliefs or desires, even when those responses are objectively incorrect or misleading. While sycophantic behavior is typically studied through explicit prompt perturbations or behavioral benchmarks, latent sycophantic patterns may persist hidden within a model’s internal representations and outp...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e5b1accf-579b-53a0-b101-ab3864d3141c
- DOI
- 10.14293/pr2199.003532.v1
