Back to search

Article

Meta-PPO: Lightweight Meta-Control for Online Joint Hyperparameter Adaptation in Proximal Policy Optimization

2026-07-15

Abstract excerpt

<title>Abstract</title> <p>Proximal Policy Optimization (PPO) is a widely used deep reinforcement-learning algorithm because its clipped surrogate objective offers a practical compromise between update stability and implementation simplicity. However, PPO remains highly sensitive to manually specified hyperparameters, particularly the learning rate, clipping range, and entropy coefficient. Fixed or open-loop sche...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
581b1e75-7538-50f7-ab2a-76a228b1ac73
DOI
10.21203/rs.3.rs-10273708/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Meta-PPO: Lightweight Meta-Control for Online Joint Hyperparameter Adaptation in Proximal Policy OptimizationDOI 10.21203/rs.3.rs-10273708/v1
Select a neighboring publication to make it the new centre.