conference · arXiv (Cornell University) · 2024

LumberChunker: Long-Form Narrative Document Segmentation

André V. Duarte, João Marques, Miguel Graça, Miguel Freire, Lei Li, Arlindo L. Oliveira · 17 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
We propose LumberChunker, a method that uses a large language model to dynamically segment long-form narrative documents into variable-sized chunks, improving retrieval performance for modern NLP tasks.
Problem
Modern NLP tasks rely on dense retrieval to access relevant context, but standard document chunking struggles to capture semantic independence because content is rarely uniform in size.
Method
We propose LumberChunker, a method that iteratively prompts a large language model to identify the exact points within sequential passages where the content begins to shift.
Results
LumberChunker outperforms the most competitive baseline by 7.37% in retrieval performance (DCG@20) and proves more effective than other chunking methods and competitive baselines like Gemini 1.5M Pro when integrated into a RAG pipeline.
Contributions
We introduce LumberChunker for dynamic document segmentation and GutenQA, a benchmark featuring 3,000 "needle in a haystack" question-answer pairs derived from 100 public domain narrative books on Project Gutenberg.
Limitations
Not specified in the abstract.
Takeaways
Allowing document segments to vary in size to better capture semantic independence significantly improves dense retrieval and RAG pipeline effectiveness.
Applications
Retrieval-augmented generation (RAG) pipelines, document search systems, and information retrieval applications.
Topics
Long-form narrative document segmentation, dense retrieval, large language models, RAG pipelines
For industry
Information technology, search engines, and artificial intelligence development
Why it matters
Advances text retrieval and RAG pipeline capabilities for processing long-form documents.

Abstract

Modern NLP tasks increasingly rely on dense retrieval methods to access up-to-date and relevant contextual information.We are motivated by the premise that retrieval benefits from segments that can vary in size such that a content's semantic independence is better captured.We propose LumberChunker, a method leveraging an LLM to dynamically segment documents, which iteratively prompts the LLM to identify the point within a group of sequential passages where the content begins to shift.To evaluate our method, we introduce GutenQA, a benchmark with 3000 "needle in a haystack" type of question-answer pairs derived from 100 public domain narrative books available on Project Gutenberg 1 .Our experiments show that Lum-berChunker not only outperforms the most competitive baseline by 7.37% in retrieval performance (DCG@20) but also that, when integrated into a RAG pipeline, LumberChunker proves to be more effective than other chunking methods and competitive baselines, such as the Gemini 1.5M Pro.

← All publications