如何解决箱线图(boxplot)中unhashable type: 'numpy.ndarray'错误?
解决seaborn.boxplot报错TypeError: unhashable type: 'numpy.ndarray'的方案
错误原因
- 传递给
sns.boxplot的f1_scores是嵌套数组结构(列表中每个元素是10个交叉验证分数的numpy数组),而seaborn的boxplot需要一维数据对应分组标签,无法直接解析二维嵌套数据。 - 直接将
np.arange生成的数组作为参数传递时,seaborn会把它当作分组键,但numpy数组属于不可哈希类型,触发unhashable type错误。 - 你尝试的
float(np.arange(.1, .9, .1))完全无效:该np.arange生成的是包含8个元素的数组,无法直接转为单个float,只会引发新的类型错误。
解决方案
方法一:使用pandas DataFrame整理数据(推荐)
seaborn基于DataFrame的接口更直观且不易出错,先将数据整理成长格式DataFrame再传入boxplot:
import pandas as pd import numpy as np import seaborn as sns import matplotlib.pyplot as plt from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import cross_val_score # 假设你的custom_f1函数定义如下(已有则忽略) def custom_f1(cutoff): from sklearn.metrics import make_scorer, f1_score def score_func(y_true, y_pred_proba): y_pred = (y_pred_proba[:, 1] >= cutoff).astype(int) return f1_score(y_true, y_pred) return make_scorer(score_func, needs_proba=True) f1_scores = [] cutoffs = np.arange(.1, .9, .1) for cutoff in cutoffs: clf = RandomForestClassifier(15, random_state=int(cutoff * 10)) f1_scores.append(cross_val_score(clf, X, y, scoring=custom_f1(cutoff), cv=10)) # 转换为长格式DataFrame data_rows = [] for cutoff, scores in zip(cutoffs, f1_scores): for score in scores: data_rows.append({"cutoff": cutoff, "f1_score": score}) df = pd.DataFrame(data_rows) # 绘制箱线图 sns.boxplot(x="cutoff", y="f1_score", data=df) sns.despine() plt.xlabel('cutoff threshold') plt.ylabel('f1 score') plt.title('F1 score varies with different cutoff thresholds') plt.show()
方法二:扁平化数据后直接传递参数
若不想使用pandas,可将f1_scores扁平化,同时为每个分数生成对应的cutoff标签:
import numpy as np import seaborn as sns import matplotlib.pyplot as plt from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import cross_val_score # 自定义评分函数(同上) def custom_f1(cutoff): from sklearn.metrics import make_scorer, f1_score def score_func(y_true, y_pred_proba): y_pred = (y_pred_proba[:, 1] >= cutoff).astype(int) return f1_score(y_true, y_pred) return make_scorer(score_func, needs_proba=True) f1_scores = [] cutoffs = np.arange(.1, .9, .1) for cutoff in cutoffs: clf = RandomForestClassifier(15, random_state=int(cutoff * 10)) f1_scores.append(cross_val_score(clf, X, y, scoring=custom_f1(cutoff), cv=10)) # 扁平化f1分数数组 flat_f1_scores = np.concatenate(f1_scores) # 生成每个分数对应的cutoff标签(每个cutoff对应10个交叉验证分数) cutoff_labels = np.repeat(cutoffs, [len(scores) for scores in f1_scores]) # 绘制箱线图 sns.boxplot(x=cutoff_labels, y=flat_f1_scores) sns.despine() plt.xlabel('cutoff threshold') plt.ylabel('f1 score') plt.title('F1 score varies with different cutoff thresholds') plt.show()
关键说明
核心问题是数据结构不匹配:seaborn的boxplot要求分组变量(x)和数值变量(y)均为一维数组,原代码中的f1_scores是二维嵌套结构,无法被正确解析。不要尝试将多维numpy数组直接转为单个float,这是错误的解决方向。
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

