This paper contributes to the existing literature on cluster analysis presenting a new method to identify the best partition, that is, the optimal number of clusters, when the (Formula presented.) -means clustering algorithm is adopted. Clustering is an unsupervised learning technique aimed at constructing partitions such that items within the same cluster are similar, while those in different clusters exhibit distinct differences. As an exploratory analysis, the effectiveness of clustering algorithms, particularly non-hierarchical ones, depends heavily on key decisions, such as determining the number of clusters in the final partition. In the literature, various cluster validity indices are available to assist in identifying the optimal partition. However, these indices often produce conflicting recommendations, which may not provide clear guidance to researchers and practitioners. To address this gap, the Multivariate Permutation Cluster Validity (MPCV) test has been introduced to compare partitions across multiple clustering performance metrics. This approach synthesizes information from widely used clustering indices, including the Silhouette, Dunn, Calinski–Harabasz, and Davies–Bouldin indices, aiming to achieve an optimal balance between internal homogeneity and external separability. Our approach is presented and discussed using both simulated and real data, highlighting its main advantages.

A Multivariate Permutation Cluster Validity Test

Elena Barzizza
;
Nicoló Biasetton;Riccardo Ceccato;Marta Disegna
2026

Abstract

This paper contributes to the existing literature on cluster analysis presenting a new method to identify the best partition, that is, the optimal number of clusters, when the (Formula presented.) -means clustering algorithm is adopted. Clustering is an unsupervised learning technique aimed at constructing partitions such that items within the same cluster are similar, while those in different clusters exhibit distinct differences. As an exploratory analysis, the effectiveness of clustering algorithms, particularly non-hierarchical ones, depends heavily on key decisions, such as determining the number of clusters in the final partition. In the literature, various cluster validity indices are available to assist in identifying the optimal partition. However, these indices often produce conflicting recommendations, which may not provide clear guidance to researchers and practitioners. To address this gap, the Multivariate Permutation Cluster Validity (MPCV) test has been introduced to compare partitions across multiple clustering performance metrics. This approach synthesizes information from widely used clustering indices, including the Silhouette, Dunn, Calinski–Harabasz, and Davies–Bouldin indices, aiming to achieve an optimal balance between internal homogeneity and external separability. Our approach is presented and discussed using both simulated and real data, highlighting its main advantages.
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11577/3593020
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
  • OpenAlex 0
social impact