You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Kmeans聚类中指定各簇数据占比?Python实现遇阻

指定簇样本比例的聚类实现方案

标准K-Means没法直接指定各簇的样本占比——它的核心逻辑是通过最小化簇内平方和划分数据,簇的大小完全由数据点到质心的距离分布决定,不会遵循预设的比例规则。要实现你要的20%、30%、50%的簇大小分配,可以用以下两种方法:

1. 使用带大小约束的K-Means变种

scikit-learn-extra库提供了支持簇大小约束的KMeans实现,能直接指定每个簇的最小/最大样本数,强制满足比例要求。

步骤与代码示例

首先安装依赖:

pip install scikit-learn-extra

然后编写代码:

import numpy as np
from sklearn_extra.cluster import KMeans
from sklearn.datasets import make_blobs

# 生成模拟数据(替换成你的真实数据)
X, _ = make_blobs(n_samples=1000, centers=3, random_state=42)
total_samples = X.shape[0]

# 定义各簇的目标比例
target_proportions = [0.2, 0.3, 0.5]
# 计算对应样本数(确保是整数)
target_sizes = [int(p * total_samples) for p in target_proportions]

# 初始化带约束的K-Means,强制簇大小等于目标值
constrained_kmeans = KMeans(
    n_clusters=3,
    size_min=target_sizes,
    size_max=target_sizes,
    random_state=42
)
cluster_labels = constrained_kmeans.fit_predict(X)

# 验证结果
unique_clusters, counts = np.unique(cluster_labels, return_counts=True)
print("各簇实际样本数:", dict(zip(unique_clusters, counts)))

这个方法会在聚类迭代过程中自动调整质心和簇分配,同时严格遵守你设定的簇大小约束,聚类效果比手动调整更合理。

2. 手动调整普通K-Means的结果

如果不想引入额外库,可以先跑普通K-Means,再手动调整簇的样本分配,直到满足比例。这种方法简单但会牺牲部分聚类质量,因为是强制移动数据点。

代码示例

import numpy as np
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=1000, centers=3, random_state=42)
total_samples = X.shape[0]
target_proportions = [0.2, 0.3, 0.5]
target_sizes = [int(p * total_samples) for p in target_proportions]

# 先跑普通K-Means
kmeans = KMeans(n_clusters=3, random_state=42)
cluster_labels = kmeans.fit_predict(X)
centers = kmeans.cluster_centers_
current_sizes = np.bincount(cluster_labels)

# 调整簇分配
for cluster_idx in range(3):
    size_diff = target_sizes[cluster_idx] - current_sizes[cluster_idx]
    if size_diff > 0:
        # 需要从其他簇移入size_diff个点
        other_clusters = [c for c in range(3) if c != cluster_idx]
        for _ in range(size_diff):
            # 找其他簇中离当前簇质心最近的点
            min_distance = np.inf
            move_point_idx = -1
            source_cluster = -1
            for c in other_clusters:
                points_in_c = np.where(cluster_labels == c)[0]
                distances = np.linalg.norm(X[points_in_c] - centers[cluster_idx], axis=1)
                closest_idx_in_c = np.argmin(distances)
                if distances[closest_idx_in_c] < min_distance:
                    min_distance = distances[closest_idx_in_c]
                    move_point_idx = points_in_c[closest_idx_in_c]
                    source_cluster = c
            # 移动点并更新计数
            cluster_labels[move_point_idx] = cluster_idx
            current_sizes[cluster_idx] += 1
            current_sizes[source_cluster] -= 1
    elif size_diff < 0:
        # 需要移出abs(size_diff)个点到其他簇
        points_in_cluster = np.where(cluster_labels == cluster_idx)[0]
        for _ in range(-size_diff):
            # 找当前簇中离其他簇质心最近的点
            min_distance = np.inf
            move_point_idx = -1
            target_cluster = -1
            for c in other_clusters:
                distances = np.linalg.norm(X[points_in_cluster] - centers[c], axis=1)
                closest_idx_in_cluster = np.argmin(distances)
                if distances[closest_idx_in_cluster] < min_distance:
                    min_distance = distances[closest_idx_in_cluster]
                    move_point_idx = points_in_cluster[closest_idx_in_cluster]
                    target_cluster = c
            # 移动点并更新计数
            cluster_labels[move_point_idx] = target_cluster
            current_sizes[cluster_idx] -= 1
            current_sizes[target_cluster] += 1

# 验证结果
print("调整后各簇样本数:", current_sizes)

这种方法的核心是不断把“更适合”其他簇的点移动过去,尽量减少对聚类效果的破坏,但本质是强制调整,结果的簇内紧凑性会比带约束的K-Means差。

内容的提问来源于stack exchange,提问作者Ashiqur Rahman Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 01:05:18