Finished · MSc

Using large language models to interact with personal information systems

Authored by João Amoroso

Supervised by Arlindo L. Oliveira

Large language models, such as ChatGPT and GPT-4 have shown remarkable abilities to interact in natural language. However, they cannot be used to access and learn from personal data, stored in email records, note taking systems or photos and videos. The objective of this dissertation is to design a system that uses large language model as the interface for personal data, using APIs and enabling the user to query, relate and retrieve information stored in different sub-systems, such as mailboxes, Google records and note taking platforms such as Obsidian. The resulting system should be able to emulate the behavior of an intelligent assistant that has access to all stored personal data and, ultimately, to answer questions about that data in a way similar to the user that owns the data. Requisites: The student should have significant programming experience, and practical knowledge of machine learning languages and environments, such as PyTorch or TensorFlow. He/she should also have interest in developing the understanding of large language models and LLM APIs. Notes: The selected student will have access to the facilities of INESC-ID and the MLKD group ( https://mlkd.idss.inesc-id.pt/ ), including computing facilities that include four DELL PowerEdge C41402 servers, eight NVIDIA 32GB Tesla V100S and eight NVIDIA 64GB Tesla A100, among other computing servers ( https://mlkd.idss.inesc-id.pt/cluster )

← All dissertations

Using large language models to interact with personal information systems | MLKD @ INESC-ID