Back to search

Article

Negative Preference Reduction in Large Language Model Unlearning: An Experimental Approach

2024-10-16

Abstract excerpt

Rapid growth in the size and complexity of language models has led to increased concerns about the presence of harmful biases and inaccuracies embedded in model outputs. Addressing these concerns autonomously, without human supervision, has become an essential challenge for improving model safety and ethical alignment. This paper introduces a novel unlearning framework that systematically reduces negative preferen...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
ac276f45-2f9a-51be-bee2-43bde962bd3a
DOI
10.22541/au.172910500.09952481/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Negative Preference Reduction in Large Language Model Unlearning: An Experimental ApproachDOI 10.22541/au.172910500.09952481/v1
Select a neighboring publication to make it the new centre.