Back to search

Article

RLHF-Aligned Open LLMs: A Comparative Survey

2025-06-30

Abstract excerpt

We survey recent open-weight large language models (LLMs) fine-tuned via Reinforcement Learning from Human Feedback (RLHF) and related AI-assisted methods, focusing on LLaMA 2 (7B/13B chat variants), LLaMA 3 (8B, 70B), Mistral 7B, Mixtral 8×7B (Sparse-MoE), Falcon 7B-Instruct, OpenAssistant-based models, Alpaca 7B, and Zephyr 7B. Closed models (GPT-4, Claude 3) are included for reference. For each model, we descri...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
54d9b434-8bdd-5e8f-af8e-d40ba5545837
DOI
10.20944/preprints202506.2381.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
RLHF-Aligned Open LLMs: A Comparative SurveyDOI 10.20944/preprints202506.2381.v1
Select a neighboring publication to make it the new centre.