RSPCA: Random Sample Partition and Clustering Approximation for ensemble learning of big data

Luengo, Mohammad; Sultan Mahmud, Mohammad; Zheng, Hua; García-Gil, Diego; García, Salvador; Zhexue Huang, Joshua

doi:https://doi.org/10.1016/j.patcog.2024.111321

dc.contributor.author	Luengo, Mohammad
dc.contributor.author	Sultan Mahmud, Mohammad
dc.contributor.author	Zheng, Hua
dc.contributor.author	García-Gil, Diego
dc.contributor.author	García, Salvador
dc.contributor.author	Zhexue Huang, Joshua
dc.date.accessioned	2025-06-23T07:06:19Z
dc.date.available	2025-06-23T07:06:19Z
dc.date.issued	2025-05
dc.identifier.uri	https://hdl.handle.net/10481/104731
dc.description.abstract	Large-scale data clustering needs an approximate approach for improving computation efficiency and data scalability. In this paper, we propose a novel method for ensemble clustering of large-scale datasets that uses the Random Sample Partition and Clustering Approximation (RSPCA) to tackle the problems of big data computing in cluster analysis. In the RSPCA computing framework, a big dataset is first partitioned into a set of disjoint random samples, called RSP data blocks that remain distributions consistent with that of the original big dataset. In ensemble clustering, a few RSP data blocks are randomly selected, and a clustering operation is performed independently on each data block to generate the clustering result of the data block. All clustering results of selected data blocks are aggregated to the ensemble result as an approximate result of the entire big dataset. To improve the robustness of the ensemble result, the ensemble clustering process can be conducted incrementally using multiple batches of selected RSP data blocks. To improve computation efficiency, we use the I-niceDP algorithm to automatically find the number of clusters in RSP data blocks and the -means algorithm to determine more accurate cluster centroids in RSP data blocks as inputs to the ensemble process. Spectral and correlation clustering methods are used as the consensus functions to handle irregular clusters. Comprehensive experiment results on both real and synthetic datasets demonstrate that the ensemble of clustering results on a few RSP data blocks is sufficient for a good global discovery of the entire big dataset, and the new approach is computationally efficient and scalable to big data.	es_ES
dc.description.sponsorship	This research has been supported by the Key Basic Research Foundation of Shenzhen under Grant No. JCYJ20220818100205012	es_ES
dc.description.sponsorship	partially supported by Project PID2023-150070NB-I00 by MICINN/AEI	es_ES
dc.description.sponsorship	part of the I+D+i project granted by C-ING-250-UGR23 co-funded by Consejería de Universidad, Investigación e Innovación and for the European Union related to the FEDER Andalucía Program 2021-2027	es_ES
dc.language.iso	eng	es_ES
dc.rights	Attribution-NonCommercial-NoDerivatives 4.0 Internacional	*
dc.rights.uri	http://creativecommons.org/licenses/by-nc-nd/4.0/	*
dc.subject	Clustering approximation	es_ES
dc.subject	Ensemble clustering	es_ES
dc.subject	Incremental clustering	es_ES
dc.subject	Ensemble learning	es_ES
dc.title	RSPCA: Random Sample Partition and Clustering Approximation for ensemble learning of big data	es_ES
dc.type	journal article	es_ES
dc.rights.accessRights	open access	es_ES
dc.identifier.doi	https://doi.org/10.1016/j.patcog.2024.111321

Ficheros en el ítem

Nombre:: 1-s2.0-S0031320324010720-main.pdf
Tamaño:: 2.706Mb
Formato:: PDF
Descripción:: Artículo principal

Este ítem aparece en la(s) siguiente(s) colección(ones)

SCI2S - Artículos

Mostrar el registro sencillo del ítem

Excepto si se señala otra cosa, la licencia del ítem se describe como Attribution-NonCommercial-NoDerivatives 4.0 Internacional