Improving the Performance of Deep Neural Networks in Vision Tasks with Attention Mechanisms.
There is no single precise definition of "attention" for neural networks. Broadly, attention mechanisms are neural network layers that aggregate information from the entire input data. They do so according to the specific problem addressed, as it depends on the input data, such as a phrase or an image. This work will focus on the computer vision task. Attention mechanisms have gained traction in natural language processing, yet their use in computer vision has been on the rise for a few years. This usage of attention mechanisms is somewhat recent and has been advancing quickly, with new architectures published often. This thesis aims to study and compare different attention mechanisms to improve the performance in image classification tasks. Three use cases related to medical imaging will be used to ensure the benefits attention mechanisms bring to real-world scenarios. The results show that there are scenarios where attention mechanisms improve the performance on medical datasets. However, the performance increase was not as consistent as expected. The experiments also show that attention mechanisms need more data than their conventional counterparts.