Back to search

Article

Fine-grained Debiasing for Large Language Modelsvia Bias Intensity and Probability Decoupling

2026-04-06

Abstract excerpt

<title>Abstract</title> <p>Large Language Models (LLMs) have demonstrated remarkable capabilities butoften inherit and even amplify social biases present in their training data. Existingdebiasing approaches—particularly those based on human preference alignment,such as Reinforcement Learning from Human Feedback (RLHF) and DirectPreference Optimization (DPO)—typically treat bias as a binary attribute, overlooking...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
6dffe502-be67-5288-bc08-530d35b82c67
DOI
10.21203/rs.3.rs-8984881/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Fine-grained Debiasing for Large Language Modelsvia Bias Intensity and Probability DecouplingDOI 10.21203/rs.3.rs-8984881/v1
Select a neighboring publication to make it the new centre.