RSPCA: Random Sample Partition and Clustering Approximation for ensemble learning of big data

Luengo, Mohammad; Sultan Mahmud, Mohammad; Zheng, Hua; García-Gil, Diego; García, Salvador; Zhexue Huang, Joshua

doi:https://doi.org/10.1016/j.patcog.2024.111321

Artículo principal (2.706Mb)

Identificadores

URI: https://hdl.handle.net/10481/104731

DOI: https://doi.org/10.1016/j.patcog.2024.111321

Exportar

Materia

Clustering approximation

Ensemble clustering

Incremental clustering

Ensemble learning

Fecha

2025-05

Patrocinador

This research has been supported by the Key Basic Research Foundation of Shenzhen under Grant No. JCYJ20220818100205012; partially supported by Project PID2023-150070NB-I00 by MICINN/AEI; part of the I+D+i project granted by C-ING-250-UGR23 co-funded by Consejería de Universidad, Investigación e Innovación and for the European Union related to the FEDER Andalucía Program 2021-2027

Resumen

Large-scale data clustering needs an approximate approach for improving computation efficiency and data scalability. In this paper, we propose a novel method for ensemble clustering of large-scale datasets that uses the Random Sample Partition and Clustering Approximation (RSPCA) to tackle the problems of big data computing in cluster analysis. In the RSPCA computing framework, a big dataset is first partitioned into a set of disjoint random samples, called RSP data blocks that remain distributions consistent with that of the original big dataset. In ensemble clustering, a few RSP data blocks are randomly selected, and a clustering operation is performed independently on each data block to generate the clustering result of the data block. All clustering results of selected data blocks are aggregated to the ensemble result as an approximate result of the entire big dataset. To improve the robustness of the ensemble result, the ensemble clustering process can be conducted incrementally using multiple batches of selected RSP data blocks. To improve computation efficiency, we use the I-niceDP algorithm to automatically find the number of clusters in RSP data blocks and the -means algorithm to determine more accurate cluster centroids in RSP data blocks as inputs to the ensemble process. Spectral and correlation clustering methods are used as the consensus functions to handle irregular clusters. Comprehensive experiment results on both real and synthetic datasets demonstrate that the ensemble of clustering results on a few RSP data blocks is sufficient for a good global discovery of the entire big dataset, and the new approach is computationally efficient and scalable to big data.

Colecciones

SCI2S - Artículos

Excepto si se señala otra cosa, la licencia del ítem se describe como Attribution-NonCommercial-NoDerivatives 4.0 Internacional