Reasoning in Large Language Models and the Dual-Process Theory of Cognition
Large Language Models (LLMs) such as GPT-4, Claude, and Gemini represent a major leap in artificial intelligence, demonstrating capabilities in language understanding, generation, and increasingly, reasoning. However, their reasoning abilities remain a topic of active investigation, raising fundamental questions about the nature of their inferential processes, their limitations, and their relation to human cognition. One fruitful perspective to approach this challenge is through the lens of dual-process theories of reasoning, particularly the distinction between System 1 (fast, intuitive, automatic thinking) and System 2 (slow, deliberate, analytical reasoning), as developed in cognitive science since the 19th century and popularized by Daniel Kahneman.
The goal of this dissertation is to explore the reasoning behaviour of LLMs and investigate how it maps onto the System 1 / System 2 framework. The project will combine theoretical insights with empirical evaluations to better understand whether—and under what conditions—LLMs exhibit characteristics of fast, intuitive responses versus slow, structured reasoning.
The main objectives of the dissertation include
- Review the literature on LLM reasoning capabilities, including benchmarks such as logical reasoning tasks, mathematical problem solving, commonsense inference, and chain-of-thought prompting. The review will also include relevant studies in cognitive science on dual-process theories.
- Formulate a conceptual mapping between typical behaviours of LLMs and features of System 1 and System 2. For instance, short unprompted completions may align with intuitive (System 1) outputs, while multi-step chain-of-thought reasoning may reflect deliberative (System 2) processes—albeit implemented via different mechanisms.
- Design and run experiments using open LLMs (e.g., LLaMA, Mistral, or GPT-4 via API access) to test their performance across tasks designed to dissociate intuitive from analytical reasoning. This may include tasks such as syllogistic reasoning, cognitive reflection tests (CRT), and problems known to elicit System 1/System 2 divergence in humans.
- Investigate the role of prompting strategies, such as zero-shot, few-shot, and chain-of-thought prompts, in shifting the LLM's behaviour between fast, heuristic-like responses and slower, structured reasoning patterns.
- Analyse the results quantitatively and qualitatively to determine to what extent LLMs emulate dual-process characteristics, and reflect on the limitations of this analogy—e.g., the lack of internal metacognition or working memory in current LLMs.
- Discuss implications for both AI and cognitive science: what LLM performance tells us about artificial reasoning systems, and whether dual-process theories can inform the design or evaluation of next-generation AI.
This dissertation will bridge computational experimentation with cognitive theory and is well suited for students interested in the intersection of AI, psychology, and the philosophy of mind. Optional extensions may include comparisons across LLM architectures, integration with neuro-symbolic methods, or the use of LLMs as models of human reasoning behaviour in behavioural science simulations.
Requisites
The student should have significant programming experience, and practical knowledge of machine learning languages and environments, such as PyTorch or TensorFlow.
Notes: The selected student will have access to the facilities of INESC-ID and the MLKD group ( https://mlkd.idss.inesc-id.pt/ ), including computing facilities that include four DELL PowerEdge C41402 servers, eight NVIDIA 32GB Tesla V100, four NVIDIA 48GB A40 and four NVIDIA 64GB Tesla A100, among other computing servers ( https://mlkd.idss.inesc-id.pt/cluster ).