Finished · MSc

Pretraining the Vision Transformer using self-supervised methods for vision-based deep reinforcement learning

Authored by Manuel Goulão

Supervised by Arlindo L. Oliveira

View the thesis →

The Vision Transformer architecture has shown to be competitive in the computer vision (CV) space where it has dethroned convolution-based networks in several benchmarks. Nevertheless, Convolutional Neural Networks (CNN) remain the preferential architecture for the representation module in Reinforcement Learning. In this work, we study pretraining a Vision Transformer using several state-of-the-art self-supervised methods and assess data-efficiency gains from this training framework. We propose a new self-supervised learning method called TOV-VICReg that extends VICReg to better capture temporal relations between observations by adding a temporal order verification task. Furthermore, we evaluate the resultant encoders with Atari games in a sample-efficiency regime, procgen games for measuring generalization and an imitation learning task for a fast and reliable comparison of the representations. Our data-efficiency results show that the vision transformer, when pretrained with TOV-VICReg, outperforms the other self-supervised methods and the non-pretrained vision transformer but still struggles to overcome a CNN. Our generalization results show some limitations in our method when used in more visually complex games which leads to degradation of the generalization performance. Nevertheless, we were able to outperform a CNN in two of the ten Atari games where we perform a 100k steps evaluation and show a consistent data-efficiency gain in comparison to the non-pretrained vision transformer. Ultimately, we believe that such approaches in Deep Reinforcement Learning (DRL) might be the key to achieving new levels of performance as seen in natural language processing and computer vision.

← All dissertations