LumberChunker: Long-Form Narrative Document Segmentation
See where this sits in the topic map →Summary AI-generated
- TL;DR
- We propose LumberChunker, a method that uses a large language model to dynamically segment long-form narrative documents into variable-sized chunks, improving retrieval performance for modern NLP tasks.
- Problem
- Modern NLP tasks rely on dense retrieval to access relevant context, but standard document chunking struggles to capture semantic independence because content is rarely uniform in size.
- Method
- We propose LumberChunker, a method that iteratively prompts a large language model to identify the exact points within sequential passages where the content begins to shift.
- Results
- LumberChunker outperforms the most competitive baseline by 7.37% in retrieval performance (DCG@20) and proves more effective than other chunking methods and competitive baselines like Gemini 1.5M Pro when integrated into a RAG pipeline.
- Contributions
- We introduce LumberChunker for dynamic document segmentation and GutenQA, a benchmark featuring 3,000 "needle in a haystack" question-answer pairs derived from 100 public domain narrative books on Project Gutenberg.
- Limitations
- Not specified in the abstract.
- Takeaways
- Allowing document segments to vary in size to better capture semantic independence significantly improves dense retrieval and RAG pipeline effectiveness.
- Applications
- Retrieval-augmented generation (RAG) pipelines, document search systems, and information retrieval applications.
- Topics
- Long-form narrative document segmentation, dense retrieval, large language models, RAG pipelines
- For industry
- Information technology, search engines, and artificial intelligence development
- Why it matters
- Advances text retrieval and RAG pipeline capabilities for processing long-form documents.
Abstract
Modern NLP tasks increasingly rely on dense retrieval methods to access up-to-date and relevant contextual information.We are motivated by the premise that retrieval benefits from segments that can vary in size such that a content's semantic independence is better captured.We propose LumberChunker, a method leveraging an LLM to dynamically segment documents, which iteratively prompts the LLM to identify the point within a group of sequential passages where the content begins to shift.To evaluate our method, we introduce GutenQA, a benchmark with 3000 "needle in a haystack" type of question-answer pairs derived from 100 public domain narrative books available on Project Gutenberg 1 .Our experiments show that Lum-berChunker not only outperforms the most competitive baseline by 7.37% in retrieval performance (DCG@20) but also that, when integrated into a RAG pipeline, LumberChunker proves to be more effective than other chunking methods and competitive baselines, such as the Gemini 1.5M Pro.