Back to search

Article

Boost Large Language Model Performance through Self-Training with Reward Guided Tree Search

2024-07-16

Abstract excerpt

Large language models have changed the field of natural language processing by enabling sophisticated language generation tasks, yet there remains a persistent challenge in enhancing their performance through autonomous learning. Introducing Process Reward Guided Tree Search to the GPT-Neo architecture offers a novel and significant advancement by enabling the model to self-train and optimize its performance witho...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
98669e1a-e900-595f-a8a0-e905f6f9e9d7
DOI
10.22541/au.172116715.57325819/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Boost Large Language Model Performance through Self-Training with Reward Guided Tree SearchDOI 10.22541/au.172116715.57325819/v1
Select a neighboring publication to make it the new centre.