Article
Evaluating reasoning LLMs’ potential to perpetuate racial and gender disease stereotypes in healthcare
2025-08-07
Abstract excerpt
This evaluation of 36,000 clinical vignettes found that next-generation reasoning large language models, o3-mini and DeepSeek-R1, frequently perpetuate racial and gender stereotypes for common medical conditions, indicating that advancements in reasoning do not inherently improve representational fairness.
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9d294710-db98-5033-be7e-654876a4ea3f
- DOI
- 10.1101/2025.08.05.25333007
