You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn中GMM的BIC准则选最优聚类k值结果不稳定的原因咨询

问题:GMM基于BIC选择最优聚类数k时,每次运行结果不同的原因?

我需要为KMeans算法及对应数据集确定最优聚类数k。在Sklearn文档中找到基于BIC准则的高斯混合模型(GMM)选择方法,参考官方示例代码适配了自身数据集,但每次运行代码得到的最优k值均不相同,请问该现象的原因是什么?

以下是适配后的代码:

import numpy as np
import pandas as pd
import itertools
from scipy import linalg
import matplotlib.pyplot as plt
import matplotlib as mpl
from sklearn import mixture
print(__doc__)
# Number of samples per component
n_samples = 440
path = 'C:/Users/Lionel/Downloads'
file = 'Wholesale customers data.csv'
data = pd.read_csv(path + '/'+file)
X = np.array(data.iloc[:,2 :])
lowest_bic = np.infty
bic = []
n_components_range = range(1, 12)
cv_types = ['spherical', 'tied', 'diag', 'full']
for cv_type in cv_types:
    for n_components in n_components_range:
        # Fit a Gaussian mixture with EM
        gmm = mixture.GaussianMixture(n_components=n_components, covariance_type=cv_type)
        gmm.fit(X)
        bic.append(gmm.bic(X))
        if bic[-1] < lowest_bic:
            lowest_bic = bic[-1]
            best_gmm = gmm
bic = np.array(bic)
color_iter = itertools.cycle(['navy', 'turquoise', 'cornflowerblue', 'darkorange'])
clf = best_gmm
print(clf)
bars = []
# Plot the BIC scores
spl = plt.subplot(2, 1, 1)
#spl = plt.plot()
for i, (cv_type, color) in enumerate(zip(cv_types, color_iter)):
    xpos = np.array(n_components_range) + .2 * (i - 2)
    bars.append(plt.bar(xpos, bic[i * len(n_components_range): (i + 1) * len(n_components_range)], width=.2, color=color))
plt.xticks(n_components_range)
plt.ylim([bic.min() * 1.01 - .01 * bic.max(), bic.max()])
plt.title('BIC score per model')
xpos = np.mod(bic.argmin(), len(n_components_range)) + .65 +\
    .2 * np.floor(bic.argmin() / len(n_components_range))
plt.text(xpos, bic.min() * 0.97 + .03 * bic.max(), '*', fontsize=14)
spl.set_xlabel('Number of components')
spl.legend([b[0] for b in bars], cv_types)
# Plot the winner
splot = plt.subplot(2, 1, 2)
Y_ = clf.predict(X)
for i, (mean, cov, color) in enumerate(zip(clf.means_, clf.covariances_, color_iter)):
    v, w = linalg.eigh(cov)
    if not np.any(Y_ == i):
        continue
    plt.scatter(X[Y_ == i, 0], X[Y_ == i, 1], .8, color=color)
    # Plot an ellipse to show the Gaussian component
    angle = np.arctan2(w[0][1], w[0][0])
    angle = 180. * angle / np.pi  # convert to degrees
    v = 2. * np.sqrt(2.) * np.sqrt(v)
    ell = mpl.patches.Ellipse(mean, v[0], v[1], 180. + angle, color=color)
    ell.set_clip_box(splot.bbox)
    ell.set_alpha(.5)
    splot.add_artist(ell)
plt.xticks(())
plt.yticks(())
plt.title('Selected GMM: full model, 2 components')
plt.subplots_adjust(hspace=.35, bottom=.02)
plt.show()

原因分析与解决方案

嘿,这个问题我之前折腾GMM的时候也碰到过!其实核心原因和GMM的EM算法特性以及代码设置有关,咱们一步步拆解:

1. EM算法的随机性导致局部最优

Sklearn的GaussianMixture默认会随机初始化模型的均值、协方差(或者用kmeans初始化,但kmeans本身也受随机种子影响)。而EM算法是局部最优算法——不同的初始值可能会让算法收敛到不同的局部极小值,这就导致每次计算出的BIC值不一样,最终选出来的最优k自然也就波动了。

2. BIC值的差异可能极其细微

有时候不同k对应的BIC值差距非常小,某次运行中k=3的BIC略低,下次可能k=2的BIC反超。这本质上是你的数据集本身聚类边界不清晰,或者不同组件数的GMM拟合效果接近,BIC准则没办法给出绝对明确的最优解。

针对性解决办法

  • 固定随机种子:在创建GaussianMixture实例时加上random_state参数,比如:
    gmm = mixture.GaussianMixture(n_components=n_components, covariance_type=cv_type, random_state=42)
    
    这样每次初始化的结果完全一致,EM收敛的结果也会固定,最优k自然就不会变了。
  • 多次运行取稳定结果:如果不想固定种子,可以循环运行多次代码,统计各个k被选为最优的次数,选出现频率最高的那个k,这样结果更可靠。
  • 预处理数据集:你用的Wholesale Customers数据特征量纲差异很大(比如Fresh和Milk的数值范围差很多),GMM对特征尺度非常敏感,建议先做标准化处理,比如:
    from sklearn.preprocessing import StandardScaler
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    
    用标准化后的X_scaled去拟合GMM,不仅拟合效果会更稳定,BIC的计算也会更合理。

内容的提问来源于stack exchange,提问作者Krukiou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:52:25