热力图低数值颜色过深是否正常?混淆矩阵可视化疑问
问题:混淆矩阵热力图低数值颜色异常偏深,是否正常?
绘制混淆矩阵热力图时发现,部分较低数值(如图中第一列的0.44%)颜色比预期更深。以下是实现代码、打印输出及结果图,请问该现象是否正常?
实现代码
categories = list(dictionary.label2idx.keys()) fig, ax = plt.subplots(figsize=(10, 10)) target_tokens_flat = np.array([label for sentence_labels in target_tokens for label in sentence_labels]) output_tokens_flat = np.array([label for sentence_labels in output_tokens for label in sentence_labels]) print(target_tokens_flat) cm = confusion_matrix(target_tokens_flat, output_tokens_flat, labels=categories) print(cm) fmt = '.2%' total_predictions = np.sum(cm) annot = np.empty_like(cm).astype(str) cm = np.array(cm) column_sums = [np.sum(cm[:, i], axis=0) for i in range(cm.shape[0])] print(column_sums) result = [(cm[:, i] / (column_sums[i]+0.0000001)).round(4) for i in range(cm.shape[1])] result = np.transpose(result) print(result) sns.heatmap(cm, annot=result, fmt=fmt, cmap='flare', xticklabels=categories, yticklabels=categories,vmin=0, vmax=100)
结果图

打印输出
['O' 'O' 'O' ... 'O' 'O' 'O'] [[31041 522 208 206 26 52 1 12 6 0 0] [ 395 1284 41 20 3 11 0 0 0 0 0] [ 140 47 796 9 8 2 0 0 0 0 0] [ 111 9 12 367 3 0 0 0 0 0 0] [ 33 14 6 2 138 4 0 0 0 0 0] [ 43 28 7 2 3 132 0 0 0 0 0] [ 3 0 0 0 0 0 33 3 0 0 0] [ 16 5 0 0 0 0 2 11 0 0 0] [ 10 0 0 0 0 0 0 1 41 0 0 0] [ 0 0 0 0 0 0 0 0 0 0 0] [ 1 0 0 0 0 0 0 1 0 0 0]] [31793, 1909, 1070, 606, 181, 201, 36, 28, 47, 0, 0] [[9.763e-01 2.734e-01 1.944e-01 3.399e-01 1.436e-01 2.587e-01 2.780e-02 4.286e-01 1.277e-01 0.000e+00 0.000e+00] [1.240e-02 6.726e-01 3.830e-02 3.300e-02 1.660e-02 5.470e-02 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00] .... [3.000e-04 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 3.570e-02 8.723e-01 0.000e+00 0.000e+00] [0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00] [0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 0.000e+00 3.570e-02 0.000e+00 0.000e+00 0.000e+00]]
回答
这个现象不正常,核心问题是你把热力图的颜色映射数据和标注数据搞混了:
- 传给
sns.heatmap的第一个参数是原始混淆矩阵cm(计数数据,比如第一列第一个值是31041),颜色是根据这些大数的范围来映射的; - 但你显示的标注是归一化后的百分比
result(比如第一列某值是0.44%),两者完全不匹配,导致视觉上出现“低数值颜色深”的错觉。
举个例子:第一列的0.44%对应的原始计数是10,这个数值在原始cm里虽然小,但和其他列的极小值(比如0)比起来,会被赋予较深的颜色,但标注的是百分比,所以看起来颜色和数值不对应。
修正方案
把热力图的数据源换成归一化后的result,同时调整vmin和vmax匹配数据范围,修改后的关键代码如下:
# 方案1:用0-1的比例数据,调整vmin/vmax sns.heatmap(result, annot=result, fmt=fmt, cmap='flare', xticklabels=categories, yticklabels=categories, vmin=0, vmax=1) # 方案2:转成百分比后使用 result_percent = result * 100 sns.heatmap(result_percent, annot=result_percent, fmt='.2f%', cmap='flare', xticklabels=categories, yticklabels=categories, vmin=0, vmax=100)
这样颜色就会和标注的百分比数值对应,不会再出现颜色和数值不符的情况。
内容的提问来源于stack exchange,提问作者feedback
相关产品推荐
相关产品推荐

