Article
Negative Preference Reduction in Large Language Model Unlearning: An Experimental Approach
2024-10-16
Abstract excerpt
Rapid growth in the size and complexity of language models has led to increased concerns about the presence of harmful biases and inaccuracies embedded in model outputs. Addressing these concerns autonomously, without human supervision, has become an essential challenge for improving model safety and ethical alignment. This paper introduces a novel unlearning framework that systematically reduces negative preferen...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- ac276f45-2f9a-51be-bee2-43bde962bd3a
- DOI
- 10.22541/au.172910500.09952481/v1
