Article
Learning Safe Behaviour via Justified Human Preferences and Hypothetical Queries
2022-12-29
Abstract excerpt
<title>Abstract</title> <p>Although reinforcement learning is a powerful paradigm for agent sequential decision-making, it cannot be used in its traditional form in most safety-critical environments. Human feedback can enable an agent to learn a good policy while avoiding unsafe states, but at the cost of human time. We present JPAL-HA, a model for safe learning in safety-critical environments that is grounded on...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 0f3305e2-2baf-5eed-a3cb-ec516323a750
- DOI
- 10.21203/rs.3.rs-2406802/v1
