DE-COP: Detecting Copyrighted Content in Language Models Training Data
See where this sits in the topic map →Summary AI-generated
- TL;DR
- We propose DE-COP, a method to determine whether copyrighted content was included in the undisclosed training data of a language model.
- Problem
- Language models are typically trained on undisclosed data, making it difficult to detect if copyrighted content was used in the training process.
- Method
- DE-COP probes a language model using multiple-choice questions that feature both verbatim text excerpts and their paraphrases. The approach is evaluated using BookTection, a new benchmark of excerpts from 165 books published before and after a model's training cutoff.
- Results
- DE-COP surpasses the prior best method by 9.6% in detection performance (AUC) on models with logits available. Furthermore, it achieves an average accuracy of 72% for detecting suspect books on fully black-box models, where prior methods yield approximately 4% accuracy.
- Contributions
- The introduction of the DE-COP detection method and the BookTection benchmark, along with open-source code and datasets.
- Limitations
- Not specified in the abstract.
- Takeaways
- The code and datasets are publicly available at https://github.com/LeiLiLab/DE-COP to support further research in detecting copyrighted training data.
- Applications
- Not specified in the abstract.
- Topics
- Detecting Copyrighted Content; Language Models Training Data
- For industry
- Language technology and artificial intelligence development.
- Why it matters
- Advances our ability to audit and verify the training data used by language models.
Abstract
How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed? We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a method to determine whether a piece of copyrighted content was included in training. DE-COP's core approach is to probe an LLM with multiple-choice questions, whose options include both verbatim text and their paraphrases. We construct BookTection, a benchmark with excerpts from 165 books published prior and subsequent to a model's training cutoff, along with their paraphrases. Our experiments show that DE-COP surpasses the prior best method by 9.6% in detection performance (AUC) on models with logits available. Moreover, DE-COP also achieves an average accuracy of 72% for detecting suspect books on fully black-box models where prior methods give approximately 4% accuracy. The code and datasets are available at https://github.com/LeiLiLab/DE-COP.