如何使用两个DataFrame绘制带双图例的聚类结果3D散点图
问题:基于两个DataFrame绘制带双图例的3D聚类散点图报错
最近我在尝试使用2个不同的DataFrame绘制3D散点图,目标是输出包含2个图例的3D散点图用于展示聚类算法的运行结果。现有主DataFrame df1 包含3个特征,结构如下:
+-----+------------+----------+----------+ | id| x| y| z| +-----+------------+----------+----------+ | row0| -6.0776997|-2.9096103|-1.5181729| | row1| -1.0122601| 7.322841|-5.4424076| | row2| -8.297007| 6.3228936| 1.1672047| | row3| -3.5071216| 4.784812|-5.4449472| | row4| -5.122823|-3.3220499|-0.5069805| | row5| -2.4764006| 8.255791| 4.409478| | row6| 7.3153954| -5.079449| -7.291215| | row7| -2.0167463| 9.303454| 7.095179| | row8| -0.2338185| -4.892681| 2.1228876| | row9| 6.565442| -6.855994|-6.7983212| |row10| -5.6902847|-6.4827404|-0.9246967| |row11|-0.017986143| 2.7632365| -8.814824| |row12| -6.9042625|-6.1491723|-3.5354295| |row13| -10.389865| 9.537853| 0.674591| |row14| 3.9688683|-6.0467844| -5.462389| |row15| -7.337052|-3.7689247| -5.261122| |row16| -8.991589| 8.738728| 3.864116| |row17| -0.18098584| 5.482743| -4.900118| |row18| 3.3193955|-6.3573766| -6.978025| |row19| -2.0266335|-3.4171724|0.48218703| +-----+------------+----------+----------+
我从聚类算法输出结果中得到了另一个DataFrame df2,生成逻辑和结构如下:
print("==========================Short report==================================== ") n_clusters = model.summary.k #n_clusters print("Number of predicted clusters: " + str(n_clusters)) cluster_Sizes = model.summary.clusterSizes #cluster_Sizes col = ['size'] df2 = pd.DataFrame(cluster_Sizes, columns=col).sort_values(by=['size'], ascending=True) #sorting cluster_Sizes = df2["size"].unique() print("Size of predicted clusters: " + str(cluster_Sizes)) clusterSizes #==========================Short report==================================== #Number of predicted clusters: 10 #Size of predicted clusters: [ 486 496 504 529 985 998 999 1003 2000] +-----+----------+ | |prediction| +-----+----------+ | 2| 486| | 6| 496| | 0| 504| | 8| 529| | 5| 985| | 9| 998| | 7| 999| | 3| 1003| | 1| 2000| | 4| 2000| +-----+----------+
df2的索引列是预测得到的簇标签,我已经将预测簇标签关联到主DataFrame中,但未关联簇大小,关联后的df1结构如下:
+-----+----------+------------+----------+----------+ | id|prediction| x| y| z| +-----+----------+------------+----------+----------+ | row0| 9| -6.0776997|-2.9096103|-1.5181729| | row1| 4| -1.0122601| 7.322841|-5.4424076| | row2| 1| -8.297007| 6.3228936| 1.1672047| | row3| 4| -3.5071216| 4.784812|-5.4449472| | row4| 3| -5.122823|-3.3220499|-0.5069805| | row5| 1| -2.4764006| 8.255791| 4.409478| | row6| 5| 7.3153954| -5.079449| -7.291215| | row7| 1| -2.0167463| 9.303454| 7.095179| | row8| 7| -0.2338185| -4.892681| 2.1228876| | row9| 5| 6.565442| -6.855994|-6.7983212| |row10| 3| -5.6902847|-6.4827404|-0.9246967| |row11| 4|-0.017986143| 2.7632365| -8.814824| |row12| 9| -6.9042625|-6.1491723|-3.5354295| |row13| 1| -10.389865| 9.537853| 0.674591| |row14| 2| 3.9688683|-6.0467844| -5.462389| |row15| 9| -7.337052|-3.7689247| -5.261122| |row16| 1| -8.991589| 8.738728| 3.864116| |row17| 4| -0.18098584| 5.482743| -4.900118| |row18| 2| 3.3193955|-6.3573766| -6.978025| |row19| 7| -2.0266335|-3.4171724|0.48218703| +-----+----------+------------+----------+----------+
现在我希望通过如下函数绘制3D散点图并配置两个独立图例:
color_names = ["red", "blue", "yellow", "black", "pink", "purple", "orange"] def plot_3d_transformed_data(df, title, colors="red"): # Imports. import matplotlib as mpl import matplotlib.pyplot as plt from mpl_toolkits.mplot3d import Axes3D import pandas as pd import numpy as np import plotly.express as px import matplotlib.cm as cm # Figure. figure = plt.figure(figsize=(12, 10)) ax = figure.add_subplot(projection="3d") ax.set_xlabel("PC1: x") ax.set_ylabel("PC2: y") ax.set_zlabel("PC3: z") ax.set_title("scatter 3D legend") # Data and 3D scatter. #colors = ["red", "blue", "yellow", "black", "pink", "purple", "orange", "black", "red" ,"blue"] colors = cm.rainbow(np.linspace(0, 1, len(cluster_Sizes))) # Create your plot #px.scatter(df1, x='x', y='y', size=df2['size'], color='jet') sc = ax.scatter(df1.x, df1.y, df1.z, alpha=0.6, c=colors, sizes=df2['size'], marker="o") # Legend 1. handles, labels = sc.legend_elements(prop="sizes", alpha=0.6) legend1 = ax.legend(handles, labels, bbox_to_anchor=(1, 1), loc="upper right", title="Sizes") ax.add_artist(legend1) # <- this is important. # Legend 2. unique_colors = set(colors) handles = [] labels = [] for n, color in enumerate(unique_colors, start=1): artist = mpl.lines.Line2D([], [], color=color, lw=0, marker="o") handles.append(artist) labels.append(str(n)) legend2 = ax.legend(handles, labels, bbox_to_anchor=(0.05, 0.05), loc="lower left", title="Classes") figure.show()
当前遇到两个报错:一是c参数长度和x/y参数长度不一致,报ValueError: 'c' argument has 9 elements, which is inconsistent with 'x' and 'y' with size 10000.;二是两个DataFrame长度不匹配导致size参数报错:ValueError: s must be a scalar, or the same size as x and y。
我知道可以将df2的size字段关联到df1中解决,但该方案开销较大不是最优解。想请教有没有更优雅的方式修改plot_3d_transformed_data()函数,实现可以同时标注预测簇标签和簇大小的可视化效果。
期望输出效果如下图所示:
解决方案
两个报错的核心原因是传入的颜色、大小数组长度与样本点总数不匹配,无需全量关联两个DataFrame,仅利用df1已有的prediction字段做映射即可,修改后的代码如下:
import matplotlib as mpl import matplotlib.pyplot as plt from mpl_toolkits.mplot3d import Axes3D import numpy as np import matplotlib.cm as cm def plot_3d_transformed_data(df1, df2, title): # 预先生成簇标签到颜色、大小的映射表 # 1. 颜色映射:每个簇标签对应唯一颜色 unique_preds = df2.index.unique() color_map = {pred: cm.rainbow(i/len(unique_preds)) for i, pred in enumerate(unique_preds)} # 2. 大小映射:每个簇标签对应对应簇大小 size_map = df2['prediction'].to_dict() # 按样本生成对应的颜色、大小数组 point_colors = df1['prediction'].map(color_map).tolist() point_sizes = df1['prediction'].map(size_map).tolist() # 绘图逻辑 figure = plt.figure(figsize=(12, 10)) ax = figure.add_subplot(projection="3d") ax.set_xlabel("PC1: x") ax.set_ylabel("PC2: y") ax.set_zlabel("PC3: z") ax.set_title(title) sc = ax.scatter(df1.x, df1.y, df1.z, alpha=0.6, c=point_colors, s=point_sizes, marker="o") # 大小图例 handles, labels = sc.legend_elements(prop="sizes", alpha=0.6) legend1 = ax.legend(handles, labels, bbox_to_anchor=(1, 1), loc="upper right", title="Sizes") ax.add_artist(legend1) # 类别图例 handles = [] labels = [] for pred, color in color_map.items(): artist = mpl.lines.Line2D([], [], color=color, lw=0, marker="o") handles.append(artist) labels.append(f"簇 {pred}") legend2 = ax.legend(handles, labels, bbox_to_anchor=(0.05, 0.05), loc="lower left", title="Classes") plt.show()
修改说明
- 用字典预先生成簇标签到颜色、大小的映射,仅遍历一次簇标签即可,时间复杂度是O(K),K是簇的数量,远小于样本量N,不会有额外的大开销
- 利用pandas的
map方法批量生成每个样本点对应的颜色、大小数组,是pandas原生的向量化操作,效率远高于手动循环匹配 - 类别图例直接基于预定义的颜色映射生成,避免了去重颜色可能出现的顺序错乱问题
- 把依赖的导入移到函数外部,避免每次调用函数重复导入,提升运行效率
内容的提问来源于stack exchange,提问作者Mario
相关产品推荐
相关产品推荐

