You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K-Means聚类如何设置每个簇的最小数据量?

K-Means能否设置每个簇的最小样本量?

原生的sklearn.cluster.KMeans不支持直接设置簇的最小样本数量,因为它的核心逻辑是通过最小化簇内平方和来划分数据,不会主动约束每个簇的样本规模,所以容易出现簇大小差异悬殊的情况,尤其是当簇总数设置得较大时(比如你这里的120簇)。

针对你的问题,有几种可行的解决思路:


1. 后处理调整小簇

训练完KMeans后,手动合并或重新分配小簇的样本:

  • 统计每个簇的样本数量,找出样本量小于阈值(比如5)的小簇;
  • 对小簇中的每个样本,计算它到其他所有簇中心的距离,将其分配到距离最近的非小簇中;
  • 重新计算簇中心,或者直接使用调整后的标签进行后续分析。

示例代码(基于你的场景):

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import numpy as np

# 先训练KMeans得到初始簇标签
kmeans = KMeans(n_clusters=120, max_iter=1500, init='k-means++', random_state=42)
cluster_labels = kmeans.fit_predict(vec_matrix_pca)
centers = kmeans.cluster_centers_

# 统计每个簇的样本数
cluster_counts = np.bincount(cluster_labels)
min_size = 5
small_clusters = np.where(cluster_counts < min_size)[0]

# 处理小簇样本
for label in small_clusters:
    # 获取当前小簇的所有样本索引
    sample_indices = np.where(cluster_labels == label)[0]
    # 计算这些样本到所有簇中心的距离
    distances = np.linalg.norm(vec_matrix_pca[sample_indices, :, None] - centers.T, axis=1)
    # 排除自身簇,找到最近的簇
    distances[:, label] = np.inf  # 屏蔽到自身簇的距离
    nearest_cluster = np.argmin(distances, axis=1)
    # 更新标签
    cluster_labels[sample_indices] = nearest_cluster

# 可选:重新计算调整后的簇中心
unique_labels = np.unique(cluster_labels)
new_centers = np.array([vec_matrix_pca[cluster_labels == l].mean(axis=0) for l in unique_labels])

2. 使用支持最小簇大小的聚类算法

如果不想做后处理,可以直接换用天然支持簇大小约束的算法:

  • HDBSCAN:基于密度的聚类算法,通过min_cluster_size参数直接指定每个簇的最小样本量,能自动过滤噪声点,避免小簇产生。
  • KMedoids(sklearn-extra):可以结合自定义的距离函数和簇大小约束,但需要额外安装scikit-learn-extra库。

HDBSCAN示例代码:

import hdbscan

# 初始化HDBSCAN,设置最小簇大小
clusterer = hdbscan.HDBSCAN(min_cluster_size=5, gen_min_span_tree=True)
cluster_labels = clusterer.fit_predict(vec_matrix_pca)

# 查看结果:-1代表噪声点,其余为簇标签
print("最终簇数量:", len(np.unique(cluster_labels)) - 1)
print("各簇样本数:", np.bincount(cluster_labels[cluster_labels != -1]))

3. 优化原KMeans代码的小问题

你的原代码存在重复训练的问题:kmeans.fit(vec_matrix_pca)之后又调用fit_predict,相当于重复跑了一遍KMeans训练,浪费计算资源。可以简化为:

from sklearn.cluster import KMeans    
from sklearn.metrics import silhouette_score

inertias = []
for i in range(2, 500):
    kmeans = KMeans(n_clusters=i, max_iter=1500, init='k-means++', random_state=42)
    cluster_labels = kmeans.fit_predict(vec_matrix_pca)
    inertias.append(kmeans.inertia_)
    silhouette_avg = silhouette_score(vec_matrix_pca, cluster_labels)
    print(f"簇数量 {i},轮廓系数:{round(silhouette_avg, 4)}")

内容的提问来源于stack exchange,提问作者taga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 15:50:22