2011 INCONCOInterpretableClusteringo

(Plant & Böhm, 2011) ⇒ Claudia Plant, and Christian Böhm. (2011). “INCONCO: Interpretable Clustering of Numerical and Categorical Objects.” In: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2011) Journal. ISBN:978-1-4503-0813-7 doi:10.1145/2020408.2020584

Subject Headings:

Notes

Cited By

Quotes

Author Keywords

Algorithms; clustering; data mining; minimum description length principle; mixed-type data

Abstract

The integrative mining of heterogeneous data and the interpretability of the data mining result are two of the most important challenges of today's data mining. It is commonly agreed in the community that, particularly in the research area of clustering, both challenges have not yet received the due attention. Only few approaches for clustering of objects with mixed-type attributes exist and those few approaches do not consider cluster-specific dependencies between numerical and categorical attributes. Likewise, only a few clustering papers address the problem of interpretability: to explain why a certain set of objects have been grouped into a cluster and what a particular cluster distinguishes from another. In this paper, we approach both challenges by constructing a relationship to the concept of data compression using the Minimum Description Length principle: a detected cluster structure is the better the more efficient it can be exploited for data compression. Following this idea, we can learn, during the run of a clustering algorithm, the optimal trade-off for attribute weights and distinguish relevant attribute dependencies from coincidental ones. We extend the efficient Cholesky decomposition to model dependencies in heterogeneous data and to ensure interpretability. Our proposed algorithm, INCONCO, successfully finds clusters in mixed type data sets, identifies the relevant attribute dependencies, and explains them using linear models and case-by-case analysis. Thereby, it outperforms existing approaches in effectiveness, as our extensive experimental evaluation demonstrates.

References

;

	Author	volume	Date Value	title	type	journal	titleUrl	doi	note	year
2011 INCONCOInterpretableClusteringo	Christian Böhm Claudia Plant			INCONCO: Interpretable Clustering of Numerical and Categorical Objects				10.1145/2020408.2020584		2011