Finished · MSc

Deep learning when data is scarce

Authored by Ana Pimenta Alves

Supervised by Arlindo Manuel Limede de Oliveira

Current deep learning models require enormous amounts of data to be trained. Recent studies by DeepMind show that even models like GPT-3, which is trained with 300 billion tokens, may still be “significantly undertrained”. Simply gathering more data to keep increasing the models’ performance is not biologically reasonable (as humans don’t need such quantities of data to learn), is not possible for some tasks (where obtaining more data is very expensive) and widens the gap between the researchers with the most resources and the rest of the community. There are several approaches that try to avoid this data requirement: few shot learning, self-supervision, using pre-trained models, and loss smoothing. The objective of this dissertation is to compare these approaches and analyze in particular their relative performance per dataset size. The selected student will have access to the facilities of INESC-ID and the MLKD group (https://mlkd.idss.inesc-id.pt/), including computing facilities that include two DELL PowerEdge C41402 servers and eight NVIDIA 32GB Tesla V100S, among other machines. The work will be developed using the Pytorch or TensorFlow programming platforms for machine learning and the Observable platform for data processing. The selected student will work within the scope of the Magellan project, and have access to the sources of data and financial resources made available by the project.

← All dissertations

Deep learning when data is scarce | MLKD @ INESC-ID