conference · 2007
Semi-supervised single-label text categorization using centroid-based classifiers
See where this sits in the topic map →Summary AI-generated
- TL;DR
- This research investigates how combining a small amount of labeled data with unlabeled data impacts the accuracy of centroid-based classifiers for single-label text categorization.
- Problem
- Not specified in the abstract.
- Method
- The study employs centroid-based text categorization using both labeled and unlabeled data.
- Results
- Centroid-based methods were chosen because they offer very fast processing speeds compared to other classification methods while maintaining an accuracy close to state-of-the-art approaches.
- Contributions
- Not specified in the abstract.
- Limitations
- Not specified in the abstract.
- Takeaways
- Efficiency is particularly crucial for handling very large domains, such as regular news feeds or the web.
- Applications
- Not specified in the abstract.
- Topics
- Not specified in the abstract.
- For industry
- Not specified in the abstract.
- Why it matters
- Not specified in the abstract.
Abstract
In this paper we study the effect of using unlabeled data in conjunction with a small portion of labeled data on the accuracy of a centroid-based classifier used to perform single-label text categorization. We chose to use centroid-based methods because they are very fast when compared with other classification methods, but still present an accuracy close to that of the state-of-the-art methods. Efficiency is particularly important for very large domains, like regular news feeds, or the web.