SSC : Statistical Subspace ClusteringReport as inadecuate

SSC : Statistical Subspace Clustering - Download this document for free, or read online. Document in PDF available to download.

1 GRAppA - LIFL - Groupe de Recherche en Apprentissage Automatique 2 MOSTRARE - Modeling Tree Structures, Machine Learning, and Information Extraction LIFL - Laboratoire d-Informatique Fondamentale de Lille, Inria Lille - Nord Europe 3 MPI - Max Planck Institute for Biological Cybernetics

Abstract : Subspace clustering is an extension of traditional clustering that seeks to find clusters in different subspaces within a dataset. This is a particularly important challenge with high dimensional data where the curse of dimensionality occurs. It has also the benefit of providing smaller descriptions of the clusters found. Existing methods only consider numerical databases and do not propose any method for clusters visualization. Besides, they require some input parameters difficult to set for the user. The aim of this paper is to propose a new subspace clustering algorithm, able to tackle databases that may contain continuous as well as discrete attributes, requiring as few user parameters as possible, and producing an interpretable output. We present a method based on the use of the well-known EM algorithm on a probabilistic model designed under some specific hypotheses, allowing us to present the result as a set of rules, each one defined with as few relevant dimensions as possible. Experiments, conducted on artificial as well as real databases, show that our algorithm gives robust results, in terms of classification and interpretability of the output.

Author: Laurent Candillier - Isabelle Tellier - Fabien Torre - Olivier Bousquet -



Related documents