pandas plot.scatter传入分类列作为c参数时触发ValueError报错
报错原因
pandas 内置的 plot.scatter 方法底层对接 matplotlib 原生散点接口,c 参数只接受三类合法输入:
- 单个合法颜色值(如
'red'、'#2ecc71') - 和数据行长度一致的颜色值序列
- 数值型序列,用于连续值的颜色映射
传入的ClusterName是存储细胞类型名的字符串分类列,不属于上述三类输入范围,因此触发ValueError。
实现方案
方案1:使用seaborn绘制(推荐,零额外转换逻辑)
seaborn 原生支持字符串分类列作为颜色映射维度,会自动完成分类配色、图例生成,不需要手动做编码转换:
import seaborn as sns import matplotlib.pyplot as plt ax = sns.scatterplot( data=sorted_df, x="ICOS - costimulator:Cyc_14_ch_4", y="PD-1 - checkpoint:Cyc_12_ch_4", hue="ClusterName", # 分类列直接传给hue参数即可 palette="viridis", s=50 ) # 调整图例位置,避免遮挡绘图区域 plt.legend(bbox_to_anchor=(1.02, 1), loc="upper left") plt.tight_layout() plt.show()
方案2:无额外依赖,手动编码分类值
如果不想引入seaborn依赖,可以先把字符串分类转换为整数编码,再自定义颜色映射和图例:
import pandas as pd import matplotlib.pyplot as plt from matplotlib.colors import ListedColormap import numpy as np # 若ClusterName已经是pandas Categorical类型,可直接用.cat.codes获取编码 if pd.api.types.is_categorical_dtype(sorted_df["ClusterName"]): code_series = sorted_df["ClusterName"].cat.codes labels = sorted_df["ClusterName"].cat.categories else: labels = sorted_df["ClusterName"].unique() label_map = {v:i for i, v in enumerate(labels)} code_series = sorted_df["ClusterName"].map(label_map) # 生成对应分类数量的viridis配色 cmap = plt.cm.get_cmap("viridis", len(labels)) custom_cmap = ListedColormap(cmap(np.arange(len(labels)))) # 传入编码后的数值序列绘图 ax = sorted_df.plot.scatter( x="ICOS - costimulator:Cyc_14_ch_4", y="PD-1 - checkpoint:Cyc_12_ch_4", c=code_series, colormap=custom_cmap, s=50 ) # 手动生成分类型图例 handles = [ plt.Line2D([0], [0], marker="o", color="w", markerfacecolor=cmap(i), markersize=10) for i in range(len(labels)) ] ax.legend(handles, labels, bbox_to_anchor=(1.02, 1), loc="upper left") plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Mohammed Zidane
相关产品推荐
相关产品推荐

