Open for application · MSc

Generation of New Scientific Knowledge using LLMs

Supervised by Arlindo L. Oliveira, Vitória Cruz

Large Language Models (LLMs) have achieved impressive capabilities, yet they remain largely constrained by the data and distributions they were trained on. Scientific discovery, however, requires generating new knowledge. This raises the question: how can we enable LLMs to move beyond their training data and produce novel scientific knowledge? The scientific method relies on experimentation, but text-only chatbot LLMs receive no feedback on the hypotheses they generate unless they are grounded in environments where their ideas can be executed and evaluated. Recent work explores agentic frameworks in which LLMs write and execute scientific code, shifting the model from a passive text generator to an active agent capable of proposing ideas, testing them, and iterating toward better solutions. Recent systems such as AlphaEvolve [1] and other “AI Scientist” [2] frameworks highlight the promise of this approach, having already discovered algorithmic improvements that sometimes surpass the state-of-the-art. So far, these frameworks have not been widely applied to AI itself, although there are early examples, such as work at Google using AlphaEvolve to improve parts of the Gemini training infrastructure [1], the Darwin Gödel Machine (a self-improving LLM-based agent) [3], and applications to kernel optimization [4] [5] More recently, Andrej Karpathy’s “autoresearch” exposes a small LLM training pipeline to coding agents. The goal of this master’s thesis is to contribute to this direction by identifying suitable problems in AI research that can be tackled with automated research frameworks and applying (and possibly improving) these automated research frameworks to solve them.

The core objectives of this thesis are

  • Review existing work on automated research frameworks and understand what kinds of AI problems they have been applied to.
  • Identify a small number of AI research problems that are simple enough to experiment with, but still interesting (for example, improving how LLMs generate answers by testing different prompting or sampling strategies).
  • Define an evaluation metric for each problem and build an experimental setup where many solutions can be tested efficiently (e.g. by using tools such as vLLM to speed up inference). Faster experiments will allow more ideas to be explored.
  • Apply an automated research framework (e.g., ShinkaEvolve [6], an open-source variation of AlphaEvolve) to these problems and analyze the results, looking at the quality, diversity, and originality of the solutions found.
  • Release the experimental setup as open-source software and write a paper with the results of the experiments.

Requisites

The student should have interest in the field of theory of mind and significant programming experience, and practical knowledge of machine learning languages and environments, such as PyTorch or TensorFlow.

Notes: The selected student will have access to the facilities of INESC-ID and the MLKD group (https://mlkd.idss.inesc-id.pt/), including computing facilities that include four DELL PowerEdge C41402 servers, eight NVIDIA 32GB Tesla V100, four NVIDIA 48GB A40 and four NVIDIA 64GB Tesla A100, among other computing servers (https://mlkd.idss.inesc-id.pt/cluster).

[1] Novikov, Alexander, et al. "Alphaevolve: A coding agent for scientific and algorithmic discovery." arXiv preprint arXiv:2506.13131 (2025).

[2] Lu, Chris, et al. "The ai scientist: Towards fully automated open-ended scientific discovery." arXiv preprint arXiv:2408.06292 (2024).

[3] Jenny Zhang, Shengran Hu, Cong Lu, Robert LangeJeff Clune, Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents, arXiv preprint

[4] Yuksekgonul, Mert, et al. "Learning to discover at test time." arXiv preprint arXiv:2601.16175 (2026).

[5] Liao, Gang, et al. "Kernelevolve: Scaling agentic kernel coding for heterogeneous ai accelerators at meta." arXiv preprint arXiv:2512.23236 (2025).

[6] Lange, Robert Tjarko, Yuki Imajuku, and Edoardo Cetin. "Shinkaevolve: Towards open-ended and sample-efficient program evolution." arXiv preprint arXiv:2509.19349 (2025).

← All dissertations