Representation-Based Robustness in Goal-Conditioned Reinforcement Learning

Published in AAAI Conference on Artificial Intelligence 2024, 2024

We study adversarial robustness in goal-conditioned reinforcement learning (GCRL), an area previously unexplored. We first demonstrate that attacks and robust representation training methods designed for traditional RL become less effective when applied to GCRL. We then introduce a Semi-Contrastive Representation attack that requires only policy function information and operates during deployment, along with Adversarial Representation Tactics that combine adversarial augmentation with regularization to strengthen agent robustness against perturbations. We validate our methods across multiple state-of-the-art GCRL algorithms and release the ReRoGCRL toolkit publicly.