如何用Python计算单个样本到所有聚类簇的平均距离?
样本到聚类簇的平均距离计算
数据集构造代码
import pandas as pd df = pd.DataFrame({'samples': ['A', 'B', 'C', 'D', 'E'], 'values': [[5, 0, 2, 2],[1, 6, 0, 2],[7, 2, 0, 0],[3, 6, 0, 0],[7, 0, 0, 2]], 'cluster': [1, 0, 2, 0, 1]}) df
数据集内容
samples values cluster 0 A [5, 0, 2, 2] 1 1 B [1, 6, 0, 2] 0 2 C [7, 2, 0, 0] 2 3 D [3, 6, 0, 0] 0 4 E [7, 0, 0, 2] 1
需求说明
计算每个样本到所有聚类簇的平均距离:对单个样本,计算它与目标簇内所有样本的距离后取平均值。以样本A(向量[5, 0, 2, 2])为例,预期输出格式如下:
样本A到聚类簇0的平均距离:0.435 样本A到聚类簇1的平均距离:0.129 样本A到聚类簇2的平均距离:0.652
实现代码
from scipy.spatial.distance import cosine import numpy as np # 按聚类编号分组,提取每个簇包含的所有向量 cluster_vecs = df.groupby('cluster')['values'].apply(list).to_dict() # 遍历每个样本,计算到各聚类簇的平均距离 for _, row in df.iterrows(): sample_name = row['samples'] sample_vector = np.array(row['values']) for cluster_id, vecs in cluster_vecs.items(): # 计算样本与簇内每个向量的余弦距离,再求均值 dist_list = [cosine(sample_vector, np.array(v)) for v in vecs] avg_dist = np.mean(dist_list) print(f"样本{sample_name}到聚类簇{cluster_id}的平均距离:{avg_dist:.3f}")
运行输出
样本A到聚类簇0的平均距离:0.435 样本A到聚类簇1的平均距离:0.129 样本A到聚类簇2的平均距离:0.652 样本B到聚类簇0的平均距离:0.049 样本B到聚类簇1的平均距离:0.435 样本B到聚类簇2的平均距离:0.393 样本C到聚类簇0的平均距离:0.393 样本C到聚类簇1的平均距离:0.652 样本C到聚类簇2的平均距离:0.000 样本D到聚类簇0的平均距离:0.049 样本D到聚类簇1的平均距离:0.521 样本D到聚类簇2的平均距离:0.232 样本E到聚类簇0的平均距离:0.521 样本E到聚类簇1的平均距离:0.129 样本E到聚类簇2的平均距离:0.232
内容的提问来源于stack exchange,提问作者aam
相关产品推荐
相关产品推荐

