All sparse PCA models are wrong, but some are useful. Part III: Model interpretation

Camacho Páez, José; Smilde, A. K.; Saccenti, E.; Westerhuis, J. A.; Bro, R.

doi:10.1016/j.chemolab.2025.105498

1-s2.0-S0169743925001832-main.pdf (2.409Mb)

Identificadores

URI: https://hdl.handle.net/10481/106469

DOI: 10.1016/j.chemolab.2025.105498

Exportar

Editorial

Elsevier

Fecha

2025-08-19

Referencia bibliográfica

Camacho, J., Smilde, A. K., Saccenti, E., Westerhuis, J. A., & Bro, R. (2025). All sparse PCA models are wrong, but some are useful. Part III: Model interpretation. Chemometrics and Intelligent Laboratory Systems: An International Journal Sponsored by the Chemometrics Society, 266(105498), 105498. https://doi.org/10.1016/j.chemolab.2025.105498

Patrocinador

MICIU/AEI/10.13039/501100011033 (PID2023-1523010B-IOO); Universidad de Granada / CBUA (Open access)

Resumen

Sparse Principal Component Analysis (sPCA) is a popular matrix factorization that combines variance maximization and sparsity with the ultimate goal of improving data interpretation. In this series of papers we show that the factorization with sPCA can be complex to interpret even when confronted with simple data. In the first paper in this series, we demonstrated that sPCA models have limitations with respect to factorizing sparse and noise-free data accurately when loadings are overlapping. In the second paper, we showed that sPCA algorithms based on deflation can generate artifacts in high order components. We also show that scores orthogonalization and the incorporation of orthonormal loadings are suitable means to avoid large artifacts. Both approaches constrain the set of possible sPCA solutions in a very similar but poorly understood way. In particular, we study in this paper the sPCA solution by Zou et al., which according to our results represent the best sPCA algorithm of those considered in the series. Here, we provide new derivations on the model equations, the computation and interpretation of the model parameters and the selection of metaparemeters in practical cases, making sPCA an even more powerful data modeling tool.

Colecciones

DTSTC - Artículos

Excepto si se señala otra cosa, la licencia del ítem se describe como Atribución 4.0 Internacional