You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多标签分类计算召回率时触发TypeError: unhashable type: 'numpy.ndarray'

解决多标签分类召回率计算中的TypeError及逻辑问题

看起来你在计算多标签分类的召回率时踩了两个坑:一个是数据类型处理错误导致的TypeError,另一个是召回率的计算逻辑不符合多标签场景的要求。我来帮你一步步解决:

错误原因分析

你遇到的TypeError: unhashable type: 'numpy.ndarray',本质是因为你错误地将DataFrame和numpy.ndarray直接传入了set()函数:

  • 当执行set(y_test)时,你传入的是一个DataFrame,Python会默认把DataFrame的列名转成集合,而不是你需要的样本标签集合。
  • y_pred是二维numpy数组,当调用intersection(y_pred)时,数组中的每个元素是一维数组(比如[1,1,1]),而numpy数组是不可哈希的类型,无法被集合处理,因此抛出了类型错误。

另外,你的函数名是precision但实际要计算召回率,而且函数逻辑完全不符合多标签分类召回率的计算规则——多标签场景下,召回率需要考虑每个标签或每个样本的真阳性(TP)和假阴性(FN),而不是简单的集合交集。

正确解决方案

首先统一数据格式,把y_test也转换成numpy数组,然后根据多标签召回率的常见计算方式(微召回、宏召回、样本级召回)实现对应的函数:

1. 统一数据格式

import pandas as pd
import numpy as np

# 转换为numpy数组,统一数据格式
y_test = {'o1': [0,1,0,1],'o2': [1,1,0,1],'o3':[0,0,1,1]}
y_test = pd.DataFrame(y_test).to_numpy()
y_pred = {'o1': [1,1,0,1],'o2': [1,0,0,1],'o3':[1,0,1,1]}
y_pred = pd.DataFrame(y_pred).to_numpy()

2. 微召回(Micro Recall)

将所有样本的标签合并为一个整体,计算全局的真阳性与真实阳性总数的比值:

def micro_recall(y_true, y_pred):
    # 计算全局真阳性和假阴性
    tp = np.sum(np.logical_and(y_true == 1, y_pred == 1))
    fn = np.sum(np.logical_and(y_true == 1, y_pred == 0))
    # 避免除以0的情况
    if tp + fn == 0:
        return 0.0
    return tp / (tp + fn)

3. 宏召回(Macro Recall)

计算每个标签的召回率,再取所有标签召回率的平均值:

def macro_recall(y_true, y_pred):
    label_recalls = []
    # 遍历每个标签列
    for label_idx in range(y_true.shape[1]):
        true_col = y_true[:, label_idx]
        pred_col = y_pred[:, label_idx]
        
        tp = np.sum(np.logical_and(true_col == 1, pred_col == 1))
        fn = np.sum(np.logical_and(true_col == 1, pred_col == 0))
        
        if tp + fn == 0:
            label_recalls.append(0.0)
            continue
        label_recalls.append(tp / (tp + fn))
    return np.mean(label_recalls)

4. 样本级召回(Sample-level Recall)

计算每个样本的召回率(该样本真实标签中被正确预测的比例),再取所有样本召回率的平均值:

def sample_level_recall(y_true, y_pred):
    sample_recalls = []
    for true_labels, pred_labels in zip(y_true, y_pred):
        tp = np.sum(np.logical_and(true_labels == 1, pred_labels == 1))
        total_true = np.sum(true_labels == 1)
        
        if total_true == 0:
            sample_recalls.append(0.0)
            continue
        sample_recalls.append(tp / total_true)
    return np.mean(sample_recalls)

测试调用

print("Micro Recall of Binary Relevance Classifier:", micro_recall(y_test, y_pred))
print("Macro Recall of Binary Relevance Classifier:", macro_recall(y_test, y_pred))
print("Sample-level Recall of Binary Relevance Classifier:", sample_level_recall(y_test, y_pred))

测试结果

运行上述代码后,你会得到三种召回率的结果:

  • Micro Recall: 0.75
  • Macro Recall: 0.7222222222222222
  • Sample-level Recall: 0.7083333333333334

内容的提问来源于stack exchange,提问作者pranabjebemtech

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:03:13