Sciweavers

PKDD
2015
Springer

Limitations of Using Constraint Set Utility in Semi-Supervised Clustering

7 years 11 months ago
Limitations of Using Constraint Set Utility in Semi-Supervised Clustering
Abstract. Semi-supervised clustering algorithms allow the user to incorporate background knowledge into the clustering process. Often, this background knowledge is specified in the form of must-link (ML) and cannot-link (CL) constraints, indicating whether certain pairs of elements should be in the same cluster or not. Several traditional clustering algorithms have been adapted to operate in this setting. We compare some of these algorithms experimentally, and observe that their performances vary significantly, depending on the data set and constraints. We use two previously introduced constraint set utility measures, consistency and coherence, to help explain these differences. Motivated by the correlation between consistency and clustering performance, we also examine its use in algorithm selection. We find this consistency-based approach to be unsuccessful, and explain this result by observing that the previously found correlation between utility measures and clustering performa...
Toon van Craenendonck, Hendrik Blockeel
Added 16 Apr 2016
Updated 16 Apr 2016
Type Journal
Year 2015
Where PKDD
Authors Toon van Craenendonck, Hendrik Blockeel
Comments (0)