[1]TANG Yiming,LIU Zilong,GAO Jianwei.Global data-driven fuzzy cluster validity index[J].CAAI Transactions on Intelligent Systems,2026,21(3):598-616.[doi:10.11992/tis.202507010]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 3
Page number:
598-616
Column:
学术论文—机器学习
Public date:
2026-05-05
- Title:
-
Global data-driven fuzzy cluster validity index
- Author(s):
-
TANG Yiming; LIU Zilong; GAO Jianwei
-
School of Computer Science and Information, Hefei University of Technology, Hefei 230601, China
-
- Keywords:
-
clustering algorithms; fuzzy clustering; clustering validity index; fuzzy C-means algorithms; weighted compactness; weighted separation; robustness; noise robustness
- CLC:
-
TP181;TN99
- DOI:
-
10.11992/tis.202507010
- Abstract:
-
Existing cluster validity index(CVI) for fuzzy clustering often struggle to maintain accurate evaluation performance when handling datasets containing noise interference or exhibiting significant differences in cluster sizes. To address these limitations, this study proposes a global data-driven(GDD) index. The GDD index incorporates a scale-aware weighting mechanism for intra-cluster compactness and inter-cluster separation to mitigate the adverse impact of imbalanced cluster sizes on validity assessment. First, the GDD index adopts a ratio-based formulation of intra-cluster compactness to inter-cluster separation. This design prevents the undesirable monotonic increase of the index value as the number of clusters grows. Second, to obtain a more objective measure of intra-cluster compactness, the index computes the average distance among data points within each cluster, incorporating fuzzy membership degrees. Crucially, recognizing that clusters of different scales contribute unequally to the overall dataset structure, larger clusters are assigned higher weights. Specifically, the ratio of the number of samples in each cluster to the total number of samples is introduced into the compactness calculation. This weighting scheme directly influences each cluster’s contribution to the overall compactness, thereby enhancing representational fairness and accuracy. Third, for inter-cluster separation, the index comprehensively considers both the centroid distribution and the relative sizes of different clusters. Rather than treating all centroids equally, the index assigns higher weights to centroids of larger clusters. When computing the distance from each cluster centroid to the mean of all centroids, the sample-size ratio of the corresponding cluster is incorporated. This adjustment anchors the contribution of each centroid’s distance to the total separation measure, resulting in a more reasonable and balanced expression of inter-cluster separation. To evaluate the effectiveness and robustness of the GDD index, extensive experiments were conducted using three representative fuzzy clustering algorithms. Experimental results demonstrate that the GDD index consistently identifies the optimal number of clusters with high accuracy, adapts well across various fuzzy clustering frameworks, and demonstrates strong robustness in challenging scenarios with noise and highly imbalanced cluster sizes. The proposed index provides a more comprehensive and reliable evaluation of fuzzy clustering algorithms in complex, noisy environments.