Back to search

Article

Finding Rare Racial Stereotypes in LLM Responses via Adversarial Anomaly Detection

2026-05-18

Abstract excerpt

Large language models (LLMs) are increasingly aligned with antiracist norms through instruction tuning and reinforcement learning from human feedback (RLHF). However, overtly racist outputs have become rare, while more subtle, context-dependent racial stereotypes and assimilationist assumptions persist — appearing infrequently, often under specific linguistic or topical triggers. These rare stereotypical responses...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
e4a8565c-25eb-58c0-9b21-43ec145fd2c4
DOI
10.14293/pr2199.003531.v2
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Finding Rare Racial Stereotypes in LLM Responses via Adversarial Anomaly DetectionDOI 10.14293/pr2199.003531.v2
Select a neighboring publication to make it the new centre.