Handling Portuguese Varieties with Pre-Trained Language Models
Automatic language identification of written texts is a well-established area of research within Natural Language Processing (NLP). State-of-the-art algorithms often rely on n-gram character models, or more recently on neural language models, to identify the correct language of texts, with good results seen for different European languages. However, distinguishing between similar language varieties with a considerable overlap in the lexicon and in frequently used linguistic expressions, such as the case of Brazilian Portuguese (PT-BR) and European Portuguese (PT-PT), still presents significant challenges.
It is important to note that accurate identification of language varieties can nowadays have important applications in terms of training large language models. In the particular case of the Portuguese language, a predominance of Brazilian Portuguese corpora online induces linguistic traces on those models, limiting their adoption outside Brazil. To address this gap and promote the creation of European Portuguese resources, this work will address:
(a) the development of training datasets focusing on the discrimination between European and Brazilian Portuguese, e.g. leveraging abundant sources of data such as OpenSubtitles; and
(b) the fine-tuning of pre-trained language models to classify textual utterances according to language variety (i.e., PT-PT and PT-BR), and for generating transliterations between varieties.
In a second stage, the project may also explore the use of European Portuguese resources, filtered through the model that is to be developed, for fine-tuning LLMs in order to better support the PT-PT variety.
Experiments will be performed on existing corpora used in previous studies (e.g., the DSL-TL corpus, used for instance in the study Enhancing Portuguese Varieties Identification with Cross-Domain Approaches ), and the main deliverable from this project will be a technical report detailing the results of a large set of comparative experiments, envisioning the publication in a conference related to the areas of natural language processing or information retrieval.
Requisites
• Commitment and availability to work on the project (e.g., not recommended for students with other professional activities);
• Interest in the application of deep learning methods to natural language processing — previous projects involving the use of Transformer-based models will be valued;
• Good commandment of English;
• Excellent grades in courses related to the topics of the project (i.e., an average grade of 17 values or higher);
• Knowledge and experience with the use of tools like Overleaf and GitHub;
• Knowledge of Python and machine learning libraries such as PyTorch and HuggingFace Transformers;
• Preference will be given to students enrolled, or interested in enrolling, in the PhD fast track programme.
Notes: The project will be supervised by Bruno Martins (DEEC/IST and INESC-ID) and Arlindo Oliveira (DEI/IST and INESC-ID). It will be developed in the context of ongoing research projects currently being executed at the Human Language Technologies (HLT) group of INESC-ID (e.g., the Center for Responsible AI — https://centerforresponsible.ai ).
Within INESC-ID, students will have access to computational resources supporting the training of large neural networks (i.e., servers with A100 GPUs), and they are expected to interact with other HLT researchers working on similar topics (e.g., Ph.D. students in the group that can act as mentors to newcomers).
Students interested in this proposal should contact Prof. Bruno Martins ( bruno.g.martins@tecnico.ulisboa.pt ) in order to schedule an interview. One student from MEIC, named Rodrigo Farate Laia, has already expressed his interest in selecting this proposal.