conference · 2006

Empirical Evaluation of Centroid-based Models for Single-label Text Categorization

Ana Cardoso-Cachopo, Arlindo L. Oliveira, Rua Alves Redol · 10 citations

See where this sits in the topic map →

Summary AI-generated

TL;DR
Centroid-based models are widely used in text categorization due to their computational simplicity, robust behavior, and strong performance.
Problem
Text categorization tasks require predicting the category of a new document given a set of training documents with known categories, but the optimal configuration of centroid-based models for these tasks remains to be fully evaluated.
Method
The paper experimentally evaluates several centroid-based models on single-label text categorization tasks, analyzing document length normalization and two different term weighting schemes.
Results
The evaluation shows that document length normalization is not always beneficial, the traditional tf-idf term weighting approach remains very effective compared to newer approaches, one specific method for calculating class centroids consistently outperforms others, and a simple centroid-based model can achieve results comparable to top-performing Support Vector Machine (SVM) models.
Contributions
An empirical evaluation of centroid-based models, term weighting schemes, and document length normalization techniques for single-label text categorization.
Limitations
Not specified in the abstract.
Takeaways
Fast and computationally simple centroid-based models can perform comparably to complex models like SVMs when configured effectively.
Applications
Not specified in the abstract.
Topics
Text Categorization; Centroid-based Models; Term Weighting; Document Length Normalization
For industry
Not specified in the abstract.
Why it matters
Not specified in the abstract.

Abstract

Centroid-based models have been used in Text Categorization because, despite their computational simplicity, they show a robust behavior and good performance. In this paper we experimentally evaluate several centroidbased models on single-label text categorization tasks. We also analyze document length normalization and two different term weighting schemes. We show that: (1) Document length normalization is not always the best option in a classification task. (2) The traditional tfidf term weighting approach remains very effective, even when compared to more recent approaches. (3) Despite the fact that several ways to calculate the centroid of a class in a dataset have been proposed, there is one that always outperforms the others. (4) A computationally simple and fast centroid-based model can give results similar to the top-performing SVM model. 1 Introduction and Previous Work The main goal of text categorization (TC) is to derive models for the categorization of natural language text [19]. The objective is to derive models that, given a set of training documents with known categories and a new document, which is usually called the query, will predict the query’s category. Here, we are interested in the case

References within the group

Cited by (group publications)

← All publications