Article
Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment
2024-12-02
Abstract excerpt
The alignment of large language models (LLMs) is crucial for generating helpful and harmless content. Existing approaches leverage preference-based human feedback data to learn the reward function and align the LLM with the feedback data. However, these approaches focus on modeling the reward difference between the chosen and rejected demonstrations, rather than directly modeling the true reward from each demonstr...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 45a7e228-699c-54ca-b892-400bb88ac8a5
- DOI
- 10.32388/slo0cb
