Cluster ensemble selection based on relative validity indexes

M. C. Naldi*, A. C.P.L.F. Carvalho, R. J.G.B. Campello


Publikation: Bidrag til tidsskriftTidsskriftartikelForskningpeer review


Cluster ensemble aims at producing high quality data partitions by combining a set of different partitions produced from the same data. Diversity and quality are claimed to be critical for the selection of the partitions to be combined. To enhance these characteristics, methods can be applied to evaluate and select a subset of the partitions that provide ensemble results similar or better than those based on the full set of partitions. Previous studies have shown that this selection can significantly improve the quality of the final partitions. For such, an appropriate evaluation of the candidate partitions to be combined must be performed. In this work, several methods to evaluate and select partitions are investigated, most of them based on relative clustering validity indexes. These indexes select the partitions with the highest quality to participate in the ensemble. However, each relative index can be more suitable for particular data conformations. Thus, distinct relative indexes are combined to create a final evaluation that tends to be robust to changes in the application scenario, as the majority of the combined indexes may compensate the poor performance of some individual indexes. We also investigate the impact of the diversity among partitions used for the ensemble. A comparative evaluation of results obtained from an extensive collection of experiments involving state-of-the-art methods and statistical tests is presented. Based on the obtained results, a practical design approach is proposed to support cluster ensemble selection. This approach was successfully applied to real public domain data sets.

TidsskriftData Mining and Knowledge Discovery
Udgave nummer2
Sider (fra-til)259-289
StatusUdgivet - sep. 2013
Udgivet eksterntJa

Bibliografisk note

Funding Information:
The authors acknowledge the Brazilian Research Agencies CNPq and FAPESP for


Dyk ned i forskningsemnerne om 'Cluster ensemble selection based on relative validity indexes'. Sammen danner de et unikt fingeraftryk.