conference · 2004

SEQUENTIAL PATTERN MINING WITH APPROXIMATED CONSTRAINTS

Cláudia Antunes, Arlindo L. Oliveira · 13 citations

See where this sits in the topic map →

Summary AI-generated

TL;DR
Unsupervised pattern mining in sequential data often suffers from a lack of focus caused by generating an excessive number of rules.
Problem
Except in trivial datasets, sequential pattern mining uncovers an inherently large number of rules, making it difficult to extract meaningful insights.
Method
The paper proposes using constraint approximations to guide the mining process, reducing the number of discovered patterns without missing unknown information.
Results
The authors show that existing algorithms using regular languages as constraints can be adapted with minor changes to support approximate constraints.
Contributions
Not specified in the abstract.
Limitations
Not specified in the abstract.
Takeaways
The work introduces a simple algorithm, called ε-accepts, which checks whether a sequence is approximately accepted by a given regular language.
Applications
Not specified in the abstract.
Topics
Sequential pattern mining, constraint approximations, regular languages
For industry
Not specified in the abstract.
Why it matters
Not specified in the abstract.

Abstract

The lack of focus that is a characteristic of unsupervised pattern mining in sequential data represents one of the major limitations of this approach. This lack of focus is due to the inherently large number of rules that is likely to be discovered in any but the more trivial sets of sequences. Several authors have promoted the use of constraints to reduce that number, but those constraints approximate the mining task to a hypothesis test task. In this paper, we propose the use of constraint approximations to guide the mining process, reducing the number of discovered patterns without compromising the prime goal of data mining: to discover unknown information. We show that existent algorithms, that use regular languages as constraints, can be used with minor adaptations. We propose a simple algorithm (ε-accepts) that verifies if a sequence is approximately accepted by a given regular language.

References within the group

Cited by (group publications)

← All publications