ValueError:传入值形状为(3, 3)索引暗示为(3,7) 混淆矩阵转DataFrame报错
错误含义解释
这个报错的本质是你传入DataFrame的数据维度,和你指定的行列索引长度不匹配:
metrics.confusion_matrix(y_test, y_pred)返回的是3行3列的二维数组,说明你的测试标签和预测结果里只出现了3个实际分类类别,混淆矩阵只会为实际出现的类别生成对应的维度- 你指定的列参数
list_of_categories长度是7,和数据的列数3不一致,因此触发维度不匹配报错
排查步骤
你可以先运行以下代码确认问题根源:
import numpy as np # 确认混淆矩阵的实际维度 print("混淆矩阵形状:", metrics.confusion_matrix(y_test, y_pred).shape) # 确认你自定义的类别列表长度 print("自定义类别列表长度:", len(list_of_categories)) # 确认实际出现的唯一类别数量 print("实际出现的类别数:", len(np.unique(np.concatenate([y_test, y_pred]))))
修复方案
你可以根据需求选择对应方案:
方案1:仅展示实际出现的类别
不需要保留未出现的预设类别时,直接用实际出现的类别作为索引即可:
conf_mat = metrics.confusion_matrix(y_test, y_pred) actual_cats = np.unique(np.concatenate([y_test, y_pred])) df_report = pd.DataFrame(data=conf_mat, columns=actual_cats, index=actual_cats)
方案2:补全所有预设类别的维度
需要保留你自定义的7个类别、未出现的类别对应行列填0时,用以下方法生成全维度混淆矩阵:
from sklearn.preprocessing import LabelEncoder # 把自定义类别列表做编码映射 le = LabelEncoder() le.fit(list_of_categories) # 把标签和预测结果统一映射到自定义类别的编码位置 y_test_enc = le.transform(y_test) y_pred_enc = le.transform(y_pred) # 生成和自定义类别长度匹配的混淆矩阵 conf_mat_full = metrics.confusion_matrix(y_test_enc, y_pred_enc, labels=range(len(list_of_categories))) df_report = pd.DataFrame(data=conf_mat_full, columns=list_of_categories, index=list_of_categories)
内容的提问来源于stack exchange,提问作者Bahrein GFA
相关产品推荐
相关产品推荐

