Article
Meta-PPO: Lightweight Meta-Control for Online Joint Hyperparameter Adaptation in Proximal Policy Optimization
2026-07-15
Abstract excerpt
<title>Abstract</title> <p>Proximal Policy Optimization (PPO) is a widely used deep reinforcement-learning algorithm because its clipped surrogate objective offers a practical compromise between update stability and implementation simplicity. However, PPO remains highly sensitive to manually specified hyperparameters, particularly the learning rate, clipping range, and entropy coefficient. Fixed or open-loop sche...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 581b1e75-7538-50f7-ab2a-76a228b1ac73
- DOI
- 10.21203/rs.3.rs-10273708/v1
