Article
Fine-grained Debiasing for Large Language Modelsvia Bias Intensity and Probability Decoupling
2026-04-06
Abstract excerpt
<title>Abstract</title> <p>Large Language Models (LLMs) have demonstrated remarkable capabilities butoften inherit and even amplify social biases present in their training data. Existingdebiasing approaches—particularly those based on human preference alignment,such as Reinforcement Learning from Human Feedback (RLHF) and DirectPreference Optimization (DPO)—typically treat bias as a binary attribute, overlooking...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 6dffe502-be67-5288-bc08-530d35b82c67
- DOI
- 10.21203/rs.3.rs-8984881/v1
