You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AffinityPropagation聚类随机数据时出现收敛警告,未生成预期聚类

AffinityPropagation聚类未收敛,生成0个聚类问题解决

问题重现

计划使用随机生成的3类数据进行聚类分析,运行代码后出现收敛警告,输出Estimated number of clusters: 0,无法得到预期的3个聚类及可视化结果。

运行代码

import numpy as np 
import matplotlib.pyplot as plt 
from itertools import cycle
from sklearn.cluster import KMeans 
from sklearn.datasets import make_blobs
%matplotlib inline
np.random.seed(0)
X, y = make_blobs(n_samples=5000, centers=[[4,4], [-2, -1], [2, -3]], cluster_std=0.9)
plt.scatter(X[:, 0], X[:, 1], marker='.')
from sklearn.cluster import AffinityPropagation
from sklearn import metrics
af = AffinityPropagation(preference=-50).fit(X)
cluster_centers_indices = af.cluster_centers_indices_
labels = af.labels_
n_clusters_ = len(cluster_centers_indices)
print('Estimated number of clusters: %d' % n_clusters_)
plt.figure(1)
plt.clf()
colors = cycle('bgrcmykbgrcmykbgrcmykbgrcmyk')
for k, col in zip(range(n_clusters_), colors):
   class_members = labels == k
   cluster_center = X[cluster_centers_indices[k]]
   plt.plot(X[class_members, 0], X[class_members, 1], col + '.')
   plt.plot(cluster_center[0], cluster_center[1], 'o', markerfacecolor=col,
           markeredgecolor='k', markersize=14)
   for x in X[class_members]:
       plt.plot([cluster_center[0], x[0]], [cluster_center[1], x[1]], col)
plt.title('Estimated number of clusters: %d' % n_clusters_)
plt.show()

错误输出

Estimated number of clusters: 0
/Users/alexkaram/opt/anaconda3/lib/python3.9/site-packages/sklearn/cluster/_affinity_propagation.py:250: ConvergenceWarning: Affinity propagation did not converge, this model will not have any cluster centers.
warnings.warn(

问题原因

AffinityPropagation的preference参数设置不合理。该参数控制样本成为聚类中心的倾向,值越低,样本越难被选为中心。当设置为-50时,对于当前数据集来说过低,导致所有样本都无法成为聚类中心,算法无法收敛,最终生成0个聚类。

解决办法

方法1:调整preference参数

将preference设置为更适配当前数据集的值,可通过以下两种方式:

  • 直接尝试增大参数值(比如-200):
af = AffinityPropagation(preference=-200).fit(X)
  • 基于数据相似度的中位数自动设置,更通用:
from sklearn.metrics.pairwise import euclidean_distances
similarity = -euclidean_distances(X)  # AffinityPropagation默认使用负欧氏距离作为相似度
preference = np.median(similarity)
af = AffinityPropagation(preference=preference).fit(X)

方法2:改用KMeans(已知聚类数量场景)

由于你已经明确预期聚类数为3,KMeans更适合这种场景,直接指定聚类数量即可:

import numpy as np 
import matplotlib.pyplot as plt 
from itertools import cycle
from sklearn.cluster import KMeans 
from sklearn.datasets import make_blobs
%matplotlib inline
np.random.seed(0)
X, y = make_blobs(n_samples=5000, centers=[[4,4], [-2, -1], [2, -3]], cluster_std=0.9)
plt.scatter(X[:, 0], X[:, 1], marker='.')

kmeans = KMeans(n_clusters=3, random_state=0).fit(X)
labels = kmeans.labels_
n_clusters_ = 3
cluster_centers = kmeans.cluster_centers_

plt.figure(1)
plt.clf()
colors = cycle('bgrcmyk')
for k, col in zip(range(n_clusters_), colors):
    class_members = labels == k
    plt.plot(X[class_members, 0], X[class_members, 1], col + '.')
    plt.plot(cluster_centers[k, 0], cluster_centers[k, 1], 'o', markerfacecolor=col,
             markeredgecolor='k', markersize=14)
plt.title('Estimated number of clusters: %d' % n_clusters_)
plt.show()

内容的提问来源于stack exchange,提问作者Alex Karam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 04:05:19