Python多标签分类计算召回率时触发TypeError: unhashable type: 'numpy.ndarray'
解决多标签分类召回率计算中的TypeError及逻辑问题
看起来你在计算多标签分类的召回率时踩了两个坑:一个是数据类型处理错误导致的TypeError,另一个是召回率的计算逻辑不符合多标签场景的要求。我来帮你一步步解决:
错误原因分析
你遇到的TypeError: unhashable type: 'numpy.ndarray',本质是因为你错误地将DataFrame和numpy.ndarray直接传入了set()函数:
- 当执行
set(y_test)时,你传入的是一个DataFrame,Python会默认把DataFrame的列名转成集合,而不是你需要的样本标签集合。 y_pred是二维numpy数组,当调用intersection(y_pred)时,数组中的每个元素是一维数组(比如[1,1,1]),而numpy数组是不可哈希的类型,无法被集合处理,因此抛出了类型错误。
另外,你的函数名是precision但实际要计算召回率,而且函数逻辑完全不符合多标签分类召回率的计算规则——多标签场景下,召回率需要考虑每个标签或每个样本的真阳性(TP)和假阴性(FN),而不是简单的集合交集。
正确解决方案
首先统一数据格式,把y_test也转换成numpy数组,然后根据多标签召回率的常见计算方式(微召回、宏召回、样本级召回)实现对应的函数:
1. 统一数据格式
import pandas as pd import numpy as np # 转换为numpy数组,统一数据格式 y_test = {'o1': [0,1,0,1],'o2': [1,1,0,1],'o3':[0,0,1,1]} y_test = pd.DataFrame(y_test).to_numpy() y_pred = {'o1': [1,1,0,1],'o2': [1,0,0,1],'o3':[1,0,1,1]} y_pred = pd.DataFrame(y_pred).to_numpy()
2. 微召回(Micro Recall)
将所有样本的标签合并为一个整体,计算全局的真阳性与真实阳性总数的比值:
def micro_recall(y_true, y_pred): # 计算全局真阳性和假阴性 tp = np.sum(np.logical_and(y_true == 1, y_pred == 1)) fn = np.sum(np.logical_and(y_true == 1, y_pred == 0)) # 避免除以0的情况 if tp + fn == 0: return 0.0 return tp / (tp + fn)
3. 宏召回(Macro Recall)
计算每个标签的召回率,再取所有标签召回率的平均值:
def macro_recall(y_true, y_pred): label_recalls = [] # 遍历每个标签列 for label_idx in range(y_true.shape[1]): true_col = y_true[:, label_idx] pred_col = y_pred[:, label_idx] tp = np.sum(np.logical_and(true_col == 1, pred_col == 1)) fn = np.sum(np.logical_and(true_col == 1, pred_col == 0)) if tp + fn == 0: label_recalls.append(0.0) continue label_recalls.append(tp / (tp + fn)) return np.mean(label_recalls)
4. 样本级召回(Sample-level Recall)
计算每个样本的召回率(该样本真实标签中被正确预测的比例),再取所有样本召回率的平均值:
def sample_level_recall(y_true, y_pred): sample_recalls = [] for true_labels, pred_labels in zip(y_true, y_pred): tp = np.sum(np.logical_and(true_labels == 1, pred_labels == 1)) total_true = np.sum(true_labels == 1) if total_true == 0: sample_recalls.append(0.0) continue sample_recalls.append(tp / total_true) return np.mean(sample_recalls)
测试调用
print("Micro Recall of Binary Relevance Classifier:", micro_recall(y_test, y_pred)) print("Macro Recall of Binary Relevance Classifier:", macro_recall(y_test, y_pred)) print("Sample-level Recall of Binary Relevance Classifier:", sample_level_recall(y_test, y_pred))
测试结果
运行上述代码后,你会得到三种召回率的结果:
- Micro Recall:
0.75 - Macro Recall:
0.7222222222222222 - Sample-level Recall:
0.7083333333333334
内容的提问来源于stack exchange,提问作者pranabjebemtech
相关产品推荐
相关产品推荐

