Learning Goal-Conditioned Policies Offline with Self-Supervised Reward Shaping

01/05/2023
by   Lina Mezghani, et al.
0

Developing agents that can execute multiple skills by learning from pre-collected datasets is an important problem in robotics, where online interaction with the environment is extremely time-consuming. Moreover, manually designing reward functions for every single desired skill is prohibitive. Prior works targeted these challenges by learning goal-conditioned policies from offline datasets without manually specified rewards, through hindsight relabelling. These methods suffer from the issue of sparsity of rewards, and fail at long-horizon tasks. In this work, we propose a novel self-supervised learning phase on the pre-collected dataset to understand the structure and the dynamics of the model, and shape a dense reward function for learning policies offline. We evaluate our method on three continuous control tasks, and show that our model significantly outperforms existing approaches, especially on tasks that involve long-term planning.

READ FULL TEXT

page 7

page 15

research
06/24/2022

Phasic Self-Imitative Reduction for Sparse-Reward Goal-Conditioned Reinforcement Learning

It has been a recent trend to leverage the power of supervised learning ...
research
03/20/2023

Imitating Graph-Based Planning with Goal-Conditioned Policies

Recently, graph-based planning algorithms have gained much attention to ...
research
04/15/2021

Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills

We consider the problem of learning useful robotic skills from previousl...
research
06/07/2021

XIRL: Cross-embodiment Inverse Reinforcement Learning

We investigate the visual cross-embodiment imitation setting, in which a...
research
04/13/2022

What Matters in Language Conditioned Robotic Imitation Learning

A long-standing goal in robotics is to build robots that can perform a w...
research
07/18/2020

An Open-World Simulated Environment for Developmental Robotics

As the current trend of artificial intelligence is shifting towards self...
research
05/15/2023

An Offline Time-aware Apprenticeship Learning Framework for Evolving Reward Functions

Apprenticeship learning (AL) is a process of inducing effective decision...

Please sign up or login with your details

Forgot password? Click here to reset