journal · Frontiers in artificial intelligence and applications · 2023

Pretraining the Vision Transformer Using Self-Supervised Methods for Vision Based Deep Reinforcement Learning

Manuel Goulão, Arlindo L. Oliveira · 3 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
While Vision Transformers (ViTs) have largely replaced convolutional networks in computer vision, convolutional neural networks remain standard in reinforcement learning.
Problem
Convolutional neural networks remain the preferential architecture for the representation module in reinforcement learning, even though Vision Transformers have become dominant in computer vision.
Method
We study pretraining a Vision Transformer using several state-of-the-art self-supervised methods on observations from the Atari Learning Environment. To better capture temporal relations, we propose an extension of VICReg that incorporates a temporal order verification task.
Results
All evaluated methods successfully learn useful representations, avoid representational collapse, and improve data efficiency in reinforcement learning. Furthermore, the encoder pretrained with temporal order verification yielded the best overall performance, producing richer representations, focused attention maps, and sparser representation vectors.
Contributions
We demonstrate how self-supervised pretraining can be effectively applied to Vision Transformers in reinforcement learning, highlighting the critical role of capturing the temporal dimension.
Limitations
Not specified in the abstract.
Takeaways
Exploring the temporal dimension through tasks like temporal order verification significantly enhances the quality of representations learned by Vision Transformers, leading to better-performing reinforcement learning agents.
Applications
Vision-based deep reinforcement learning agents operating in environments like Atari.
Topics
Vision Transformer, Self-Supervised Learning, Deep Reinforcement Learning, Representation Learning
For industry
Not specified in the abstract.
Why it matters
Not specified in the abstract.

Abstract

The Vision Transformer architecture has shown to be competitive in the computer vision (CV) space where it has dethroned convolution-based networks in several benchmarks. Nevertheless, convolutional neural networks (CNN) remain the preferential architecture for the representation module in reinforcement learning. In this work, we study pretraining a Vision Transformer using several state-of-the-art self-supervised methods and assess the quality of the learned representations. To show the importance of the temporal dimension in this context we propose an extension of VICReg to better capture temporal relations between observations by adding a temporal order verification task. Our results show that all methods are effective in learning useful representations and avoiding representational collapse for observations from the Atari Learning Environment (ALE) which leads to improvements in data efficiency when we evaluated in reinforcement learning (RL). Moreover, the encoder pretrained with the temporal order verification task shows the best results across all experiments, with richer representations, more focused attention maps and sparser representation vectors throughout the layers of the encoder, which shows the importance of exploring such similarity dimension. With this work, we hope to provide some insights into the representations learned by ViT during a self-supervised pretraining with observations from RL environments and to understand which properties arise in the representations that lead to the best-performing agents.

← All publications