如何绘制Kmeans输出的(10,6)维质心数组及解决索引报错问题
报错原因
- 你的
center是普通NumPy多维数组,仅支持整数、切片、布尔数组等类型的索引,不支持Pandas DataFrame专属的字符串列名索引语法center['cluster'],这是报错的核心原因。 - 额外注意你当前的
clusterNumber是float类型,如果后续需要作为簇标签使用,建议先转换为整数类型,避免后续可视化等操作出现类型不兼容问题。
多维质心平面绘制方案
6维数据无法直接映射到2维平面,可选两种常用方案实现可视化:
方案1:PCA降维后绘制2D散点图
通过PCA算法把6维特征压缩到2个主成分,再映射到平面绘制,代码如下:
import numpy as np import matplotlib.pyplot as plt from sklearn.decomposition import PCA # 加载你提供的质心数组 centers = np.array([ [19.6135, 19.8452, 19.9962, 20.1065, 20.1966, 20.2832], [26.5262, 29.6227, 31.4583, 32.7302, 33.7162, 34.6274], [13.3404, 13.268, 13.2414, 13.2246, 13.2134, 13.2087], [44.3025, 47.7419, 49.3674, 50.5635, 51.4984, 52.3669], [58.331, 63.568, 66.6059, 69.222, 71.03, 72.5983], [23.26, 25.2503, 26.5113, 27.3892, 28.0659, 28.6797], [38.6445, 42.4035, 44.3822, 45.5953, 46.591, 47.4789], [30.3485, 33.8124, 35.8269, 37.2325, 38.3075, 39.2721], [48.3545, 53.1971, 56.0548, 58.1482, 59.7034, 61.11], [34.8697, 38.4072, 40.2917, 41.5594, 42.5017, 43.3741] ]) # 将float类型的簇编号转换为整数 cluster_labels = clusterNumber.astype(int) # PCA降维到2维 pca = PCA(n_components=2) centers_2d = pca.fit_transform(centers) # 绘制散点图 plt.figure(figsize=(8,6)) scatter = plt.scatter(centers_2d[:,0], centers_2d[:,1], c=cluster_labels, cmap='tab10', s=100) # 添加簇编号标注 for i, label in enumerate(cluster_labels): plt.annotate(label, (centers_2d[i,0], centers_2d[i,1]), xytext=(5,5), textcoords='offset points') plt.legend(handles=scatter.legend_elements()[0], labels=cluster_labels, title='簇编号') plt.xlabel('PCA主成分1') plt.ylabel('PCA主成分2') plt.title('Kmeans质心降维可视化结果') plt.show()
方案2:平行坐标图直接展示6维特征
不需要降维,可直接呈现每个簇在6个维度上的数值分布,代码如下:
import pandas as pd import matplotlib.pyplot as plt from pandas.plotting import parallel_coordinates # 加载质心数组,转换为DataFrame并添加簇列 centers = np.array([ [19.6135, 19.8452, 19.9962, 20.1065, 20.1966, 20.2832], [26.5262, 29.6227, 31.4583, 32.7302, 33.7162, 34.6274], [13.3404, 13.268, 13.2414, 13.2246, 13.2134, 13.2087], [44.3025, 47.7419, 49.3674, 50.5635, 51.4984, 52.3669], [58.331, 63.568, 66.6059, 69.222, 71.03, 72.5983], [23.26, 25.2503, 26.5113, 27.3892, 28.0659, 28.6797], [38.6445, 42.4035, 44.3822, 45.5953, 46.591, 47.4789], [30.3485, 33.8124, 35.8269, 37.2325, 38.3075, 39.2721], [48.3545, 53.1971, 56.0548, 58.1482, 59.7034, 61.11], [34.8697, 38.4072, 40.2917, 41.5594, 42.5017, 43.3741] ]) cluster_labels = clusterNumber.astype(int) df = pd.DataFrame(centers, columns=[f'特征{i+1}' for i in range(6)]) df['cluster'] = cluster_labels # 绘制平行坐标图 plt.figure(figsize=(10,6)) parallel_coordinates(df, 'cluster', colormap='tab10') plt.title('质心平行坐标图(6维特征直接展示)') plt.show()
内容的提问来源于stack exchange,提问作者david
相关产品推荐
相关产品推荐

