A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments

Piergiuseppe Mallozzi; Ezequiel Castellano; Patrizio Pelliccione; Gerardo Schneider; Kenji Tei

doi:10.1109/RoSE.2019.00011

A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments
Paper i proceeding, 2019

Without prior knowledge of the environment, a software agent can learn to achieve a goal using machine learning. Model-free Reinforcement Learning (RL) can be used to make the agent explore the environment and learn to achieve its goal by trial and error. Discovering effective policies to achieve the goal in a complex environment is a major challenge for RL. Furthermore, in safety-critical applications, such as robotics, an unsafe action may cause catastrophic consequences in the agent or in the environment. In this paper, we present an approach that uses runtime monitoring to prevent the reinforcement learning agent to perform 'wrong' actions and to exploit prior knowledge to smartly explore the environment. Each monitor is de?ned by a property that we want to enforce to the agent and a context. The monitors are orchestrated by a meta-monitor that activates and deactivates them dynamically according to the context in which the agent is learning. We have evaluated our approach by training the agent in randomly generated learning environments. Our results show that our approach blocks the agent from performing dangerous and safety-critical actions in all the generated environments. Besides, our approach helps the agent to achieve its goal faster by providing feedback and shaping its reward during learning.

Reinforcement learning

LTL invariants

Reward shaping

Runtime monitoring

Författare

Piergiuseppe Mallozzi

Chalmers, Data- och informationsteknik, Software Engineering

Forskning Andra publikationer

Ezequiel Castellano

The Graduate University for Advanced Studies (SOKENDAI)

Patrizio Pelliccione

Universita degli Studi dell'Aquila

Chalmers, Data- och informationsteknik, Software Engineering

Forskning Andra publikationer

Gerardo Schneider

Chalmers, Data- och informationsteknik, Formella metoder

Forskning Andra publikationer

Kenji Tei

Waseda University

Proceedings - 2019 IEEE/ACM 2nd International Workshop on Robotics Software Engineering, RoSE 2019

5-12 8823721
978-1-7281-2249-6 (ISBN)

2nd IEEE/ACM International Workshop on Robotics Software Engineering, RoSE 2019
Montreal, Canada,

Ämneskategorier (SSIF 2011)

Lärande

Robotteknik och automation

Datavetenskap (datalogi)

DOI

10.1109/RoSE.2019.00011

Publikationsdata kopplat till DOI

Mer information

Senast uppdaterat

2024-01-03

A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments Paper i proceeding, 2019

Författare

Piergiuseppe Mallozzi

Ezequiel Castellano

Patrizio Pelliccione

Gerardo Schneider

Kenji Tei

Proceedings - 2019 IEEE/ACM 2nd International Workshop on Robotics Software Engineering, RoSE 2019

Ämneskategorier (SSIF 2011)

DOI

Mer information

Senast uppdaterat

A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments
Paper i proceeding, 2019