MASA Safe-RL
Modular library for safe RL, providing baselines for a number of constraints using a variety of algorithms.
Software tools, libraries, and research prototypes from the group, with links to repositories, documentation, and papers.
Off-policy policy optimisation using a learned behaviour policy to reduce the variance of return estimates.
Compositional shielding for safe multi-agent reinforcement learning with decentralised execution.
Policy updates for continual reinforcement learning with certified safety on previously learned tasks.
Adaptive shielding for reinforcement learning that repairs GR(1) specifications after environment assumption violations.
Modular library for safe RL, providing baselines for a number of constraints using a variety of algorithms.
Reward monitoring for reinforcement learning from specifications expressed in quantitative temporal logic.