Sklearn中GMM的BIC准则选最优聚类k值结果不稳定的原因咨询
问题:GMM基于BIC选择最优聚类数k时,每次运行结果不同的原因?
我需要为KMeans算法及对应数据集确定最优聚类数k。在Sklearn文档中找到基于BIC准则的高斯混合模型(GMM)选择方法,参考官方示例代码适配了自身数据集,但每次运行代码得到的最优k值均不相同,请问该现象的原因是什么?
以下是适配后的代码:
import numpy as np import pandas as pd import itertools from scipy import linalg import matplotlib.pyplot as plt import matplotlib as mpl from sklearn import mixture print(__doc__) # Number of samples per component n_samples = 440 path = 'C:/Users/Lionel/Downloads' file = 'Wholesale customers data.csv' data = pd.read_csv(path + '/'+file) X = np.array(data.iloc[:,2 :]) lowest_bic = np.infty bic = [] n_components_range = range(1, 12) cv_types = ['spherical', 'tied', 'diag', 'full'] for cv_type in cv_types: for n_components in n_components_range: # Fit a Gaussian mixture with EM gmm = mixture.GaussianMixture(n_components=n_components, covariance_type=cv_type) gmm.fit(X) bic.append(gmm.bic(X)) if bic[-1] < lowest_bic: lowest_bic = bic[-1] best_gmm = gmm bic = np.array(bic) color_iter = itertools.cycle(['navy', 'turquoise', 'cornflowerblue', 'darkorange']) clf = best_gmm print(clf) bars = [] # Plot the BIC scores spl = plt.subplot(2, 1, 1) #spl = plt.plot() for i, (cv_type, color) in enumerate(zip(cv_types, color_iter)): xpos = np.array(n_components_range) + .2 * (i - 2) bars.append(plt.bar(xpos, bic[i * len(n_components_range): (i + 1) * len(n_components_range)], width=.2, color=color)) plt.xticks(n_components_range) plt.ylim([bic.min() * 1.01 - .01 * bic.max(), bic.max()]) plt.title('BIC score per model') xpos = np.mod(bic.argmin(), len(n_components_range)) + .65 +\ .2 * np.floor(bic.argmin() / len(n_components_range)) plt.text(xpos, bic.min() * 0.97 + .03 * bic.max(), '*', fontsize=14) spl.set_xlabel('Number of components') spl.legend([b[0] for b in bars], cv_types) # Plot the winner splot = plt.subplot(2, 1, 2) Y_ = clf.predict(X) for i, (mean, cov, color) in enumerate(zip(clf.means_, clf.covariances_, color_iter)): v, w = linalg.eigh(cov) if not np.any(Y_ == i): continue plt.scatter(X[Y_ == i, 0], X[Y_ == i, 1], .8, color=color) # Plot an ellipse to show the Gaussian component angle = np.arctan2(w[0][1], w[0][0]) angle = 180. * angle / np.pi # convert to degrees v = 2. * np.sqrt(2.) * np.sqrt(v) ell = mpl.patches.Ellipse(mean, v[0], v[1], 180. + angle, color=color) ell.set_clip_box(splot.bbox) ell.set_alpha(.5) splot.add_artist(ell) plt.xticks(()) plt.yticks(()) plt.title('Selected GMM: full model, 2 components') plt.subplots_adjust(hspace=.35, bottom=.02) plt.show()
原因分析与解决方案
嘿,这个问题我之前折腾GMM的时候也碰到过!其实核心原因和GMM的EM算法特性以及代码设置有关,咱们一步步拆解:
1. EM算法的随机性导致局部最优
Sklearn的GaussianMixture默认会随机初始化模型的均值、协方差(或者用kmeans初始化,但kmeans本身也受随机种子影响)。而EM算法是局部最优算法——不同的初始值可能会让算法收敛到不同的局部极小值,这就导致每次计算出的BIC值不一样,最终选出来的最优k自然也就波动了。
2. BIC值的差异可能极其细微
有时候不同k对应的BIC值差距非常小,某次运行中k=3的BIC略低,下次可能k=2的BIC反超。这本质上是你的数据集本身聚类边界不清晰,或者不同组件数的GMM拟合效果接近,BIC准则没办法给出绝对明确的最优解。
针对性解决办法
- 固定随机种子:在创建
GaussianMixture实例时加上random_state参数,比如:
这样每次初始化的结果完全一致,EM收敛的结果也会固定,最优k自然就不会变了。gmm = mixture.GaussianMixture(n_components=n_components, covariance_type=cv_type, random_state=42) - 多次运行取稳定结果:如果不想固定种子,可以循环运行多次代码,统计各个k被选为最优的次数,选出现频率最高的那个k,这样结果更可靠。
- 预处理数据集:你用的Wholesale Customers数据特征量纲差异很大(比如Fresh和Milk的数值范围差很多),GMM对特征尺度非常敏感,建议先做标准化处理,比如:
用标准化后的from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_scaled = scaler.fit_transform(X)X_scaled去拟合GMM,不仅拟合效果会更稳定,BIC的计算也会更合理。
内容的提问来源于stack exchange,提问作者Krukiou
相关产品推荐
相关产品推荐

