You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决箱线图(boxplot)中unhashable type: 'numpy.ndarray'错误?

解决seaborn.boxplot报错TypeError: unhashable type: 'numpy.ndarray'的方案

错误原因

  • 传递给sns.boxplot的f1_scores是嵌套数组结构(列表中每个元素是10个交叉验证分数的numpy数组),而seaborn的boxplot需要一维数据对应分组标签,无法直接解析二维嵌套数据。
  • 直接将np.arange生成的数组作为参数传递时,seaborn会把它当作分组键,但numpy数组属于不可哈希类型,触发unhashable type错误。
  • 你尝试的float(np.arange(.1, .9, .1))完全无效:该np.arange生成的是包含8个元素的数组,无法直接转为单个float,只会引发新的类型错误。

解决方案

方法一:使用pandas DataFrame整理数据(推荐)

seaborn基于DataFrame的接口更直观且不易出错,先将数据整理成长格式DataFrame再传入boxplot:

import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

# 假设你的custom_f1函数定义如下(已有则忽略)
def custom_f1(cutoff):
    from sklearn.metrics import make_scorer, f1_score
    def score_func(y_true, y_pred_proba):
        y_pred = (y_pred_proba[:, 1] >= cutoff).astype(int)
        return f1_score(y_true, y_pred)
    return make_scorer(score_func, needs_proba=True)

f1_scores = []
cutoffs = np.arange(.1, .9, .1)

for cutoff in cutoffs:
    clf = RandomForestClassifier(15, random_state=int(cutoff * 10))
    f1_scores.append(cross_val_score(clf, X, y, scoring=custom_f1(cutoff), cv=10))

# 转换为长格式DataFrame
data_rows = []
for cutoff, scores in zip(cutoffs, f1_scores):
    for score in scores:
        data_rows.append({"cutoff": cutoff, "f1_score": score})
df = pd.DataFrame(data_rows)

# 绘制箱线图
sns.boxplot(x="cutoff", y="f1_score", data=df)
sns.despine()
plt.xlabel('cutoff threshold')
plt.ylabel('f1 score')
plt.title('F1 score varies with different cutoff thresholds')
plt.show()

方法二:扁平化数据后直接传递参数

若不想使用pandas,可将f1_scores扁平化,同时为每个分数生成对应的cutoff标签:

import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

# 自定义评分函数(同上)
def custom_f1(cutoff):
    from sklearn.metrics import make_scorer, f1_score
    def score_func(y_true, y_pred_proba):
        y_pred = (y_pred_proba[:, 1] >= cutoff).astype(int)
        return f1_score(y_true, y_pred)
    return make_scorer(score_func, needs_proba=True)

f1_scores = []
cutoffs = np.arange(.1, .9, .1)

for cutoff in cutoffs:
    clf = RandomForestClassifier(15, random_state=int(cutoff * 10))
    f1_scores.append(cross_val_score(clf, X, y, scoring=custom_f1(cutoff), cv=10))

# 扁平化f1分数数组
flat_f1_scores = np.concatenate(f1_scores)
# 生成每个分数对应的cutoff标签(每个cutoff对应10个交叉验证分数)
cutoff_labels = np.repeat(cutoffs, [len(scores) for scores in f1_scores])

# 绘制箱线图
sns.boxplot(x=cutoff_labels, y=flat_f1_scores)
sns.despine()
plt.xlabel('cutoff threshold')
plt.ylabel('f1 score')
plt.title('F1 score varies with different cutoff thresholds')
plt.show()

关键说明

核心问题是数据结构不匹配:seaborn的boxplot要求分组变量(x)和数值变量(y)均为一维数组,原代码中的f1_scores是二维嵌套结构,无法被正确解析。不要尝试将多维numpy数组直接转为单个float,这是错误的解决方向。

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:40:37