Back to search

Article

Discovering Latent Sycophantic Patterns in LLMs via Distributional Anomaly Detection

2026-05-08

Abstract excerpt

Large language models (LLMs) often exhibit sycophancy—the tendency to generate responses that align with a user’s stated beliefs or desires, even when those responses are objectively incorrect or misleading. While sycophantic behavior is typically studied through explicit prompt perturbations or behavioral benchmarks, latent sycophantic patterns may persist hidden within a model’s internal representations and outp...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
e5b1accf-579b-53a0-b101-ab3864d3141c
DOI
10.14293/pr2199.003532.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Discovering Latent Sycophantic Patterns in LLMs via Distributional Anomaly DetectionDOI 10.14293/pr2199.003532.v1
Select a neighboring publication to make it the new centre.