Finished · MSc

Applying Deep Learning to Medical Images

Authored by Ricardo Jorge da Silva Diniz

Supervised by Arlindo L. Oliveira, Mário Alexandre Teles de Figueiredo.

View the thesis →

Deep convolutional networks have recently been embraced by the academic community as a competitive solution for visual recognition tasks. Among these networks, the fully convolutional neural networks have been gaining traction as they drop the traditional fully-connected layers of CNNs in favor of more convolutional layers. The original fully convolutional network, using layer skipping, was capable of achieving great results when provided enough samples. This architecture was extended into the U-Net which outperforms the FCNN, while being both faster and less computationally cumbersome than it. Both architectures are designed to work with 2D input images. However most medical images, such as ultrasounds and MRIs, are 3D. Built upon the underlying principles beyond the U-Net and the FCNN, the V-Net was created. It is a volumetric FCNN which introduces a new objective function, discards pooling layers in favor of more convolutional layers and performs residual propagation. V-Nets have achieved a good performance across all visual recognition tasks, being comparable to the state-of-the-art solutions while requiring a fraction of the processing time. In this thesis several variants of U-Net and V-Net are implemented to, firstly, attest to their good performance on visual segmentation tasks of medical data, and, secondly, to assess how the objective function, kernel’s receptive fields, residual propagation, activation functions and optimization method impact the model’s performance. A secondary objective of this thesis is to bridge the gap between theoretical knowledge and practical implementations by analyzing Google’s Tensorflow API, which was designed specifically for distributed computing based machine learning.

← All dissertations