Back to search

Article

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment

2024-12-02

Abstract excerpt

The alignment of large language models (LLMs) is crucial for generating helpful and harmless content. Existing approaches leverage preference-based human feedback data to learn the reward function and align the LLM with the feedback data. However, these approaches focus on modeling the reward difference between the chosen and rejected demonstrations, rather than directly modeling the true reward from each demonstr...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
45a7e228-699c-54ca-b892-400bb88ac8a5
DOI
10.32388/slo0cb
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model AlignmentDOI 10.32388/slo0cb
Select a neighboring publication to make it the new centre.