使用pyclustering执行Fuzzy C-Means聚类时如何获取聚类标签?
获取Fuzzy C-Means每个样本的聚类标签
针对你使用pyclustering的FCM算法获取样本聚类标签的需求,有两种直接的实现方式:
方法1:从聚类结果生成标签数组
fcm_instance.get_clusters()返回的是各簇包含的样本索引列表,我们可以通过遍历这些索引,为每个样本分配对应的簇编号:
import numpy as np import pandas as pd from pyclustering.cluster.center_initializer import kmeans_plusplus_initializer from pyclustering.cluster.fcm import fcm import random # 生成测试数据 coords = [(random.random()*2.0, random.random()*2.0) for _ in range(100)] dfcluster = pd.DataFrame(coords, columns = ['x','y']) sample = dfcluster.to_numpy() # 初始化聚类中心 initial_centers = kmeans_plusplus_initializer(sample, 5, kmeans_plusplus_initializer.FARTHEST_CENTER_CANDIDATE).initialize() # 执行FCM聚类 fcm_instance = fcm(sample, initial_centers) fcm_instance.process() clusters = fcm_instance.get_clusters() # 生成每个样本的聚类标签 labels = np.zeros(len(sample), dtype=int) for cluster_id, sample_indices in enumerate(clusters): for idx in sample_indices: labels[idx] = cluster_id # 将标签加入原DataFrame dfcluster['cluster_label'] = labels print(dfcluster.head())
方法2:通过隶属度矩阵生成硬标签
FCM算法的核心输出是隶属度矩阵,每个样本对应所有簇的隶属度值(范围0-1)。你可以通过get_membership()获取该矩阵,然后取每个样本隶属度最高的簇作为硬标签:
# 获取隶属度矩阵(形状:[样本数量, 簇数量]) membership_matrix = fcm_instance.get_membership() # 对每个样本取隶属度最大的簇ID作为标签 labels_from_membership = np.argmax(membership_matrix, axis=1) # 加入DataFrame dfcluster['cluster_label_from_membership'] = labels_from_membership print(dfcluster.head())
补充说明
- 两种方法生成的硬标签结果一致,
get_clusters()内部就是基于隶属度阈值或最大隶属度划分的簇。 - 如果你需要保留模糊聚类的特性,可以直接使用隶属度矩阵分析样本对各簇的归属程度;若需要明确的分类结果,上述硬标签生成方式即可满足需求。
内容的提问来源于stack exchange,提问作者I_Al-thamary
相关产品推荐
相关产品推荐

