使用ResNet50时Grad CAM输出所有图像热图一致问题求助
Grad CAM热图异常:所有图像输出一致的问题排查与解决
我使用Grad CAM分析ResNet50模型对测试图像的预测关键区域,但出现异常:所有测试图像的原始图各不相同,但对应的热图和Grad CAM叠加结果完全一致。
Grad CAM实现代码
from tensorflow.keras.models import Model import tensorflow as tf import numpy as np import cv2 class GradCAM: def __init__(self, model, classIdx, layerName=None): # 保存模型、用于计算类激活图的类别索引,以及可视化目标层 self.model = model self.classIdx = classIdx self.layerName = layerName # 如果未指定层名,自动查找目标输出层 if self.layerName is None: self.layerName = self.find_target_layer() def find_target_layer(self): # 反向遍历网络层,寻找最后一个4D输出的卷积层 for layer in reversed(self.model.layers): if len(layer.output_shape) == 4: return layer.name # 未找到4D层则无法应用GradCAM raise ValueError("Could not find 4D layer. Cannot apply GradCAM.") def compute_heatmap(self, image, eps=1e-8): # 构建梯度模型:输入为原模型输入,输出为目标卷积层输出和模型最终预测 gradModel = Model( inputs=[self.model.inputs], outputs=[self.model.get_layer(self.layerName).output, self.model.output]) # 记录自动微分操作 with tf.GradientTape() as tape: inputs = tf.cast(image, tf.float32) (convOutputs, predictions) = gradModel(inputs) loss = predictions[:, tf.argmax(predictions[0])] # 计算梯度 grads = tape.gradient(loss, convOutputs) # 计算引导梯度 castConvOutputs = tf.cast(convOutputs > 0, "float32") castGrads = tf.cast(grads > 0, "float32") guidedGrads = castConvOutputs * castGrads * grads # 去除batch维度 convOutputs = convOutputs[0] guidedGrads = guidedGrads[0] # 计算梯度均值作为权重,加权求和得到类激活图 weights = tf.reduce_mean(guidedGrads, axis=(0, 1)) cam = tf.reduce_sum(tf.multiply(weights, convOutputs), axis=-1) # 将类激活图resize到输入图像尺寸 (w, h) = (image.shape[2], image.shape[1]) heatmap = cv2.resize(cam.numpy(), (w, h)) # 归一化热图到[0,255]区间 numer = heatmap - np.min(heatmap) denom = (heatmap.max() - heatmap.min()) + eps heatmap = numer / denom heatmap = (heatmap * 255).astype("uint8") return heatmap def overlay_heatmap(self, heatmap, image, alpha=0.5, colormap=cv2.COLORMAP_VIRIDIS): # 为热图添加颜色映射并与原图叠加 heatmap = cv2.applyColorMap(heatmap, colormap) output = cv2.addWeighted(image, alpha, heatmap, 1 - alpha, 0) return (heatmap, output)
热图可视化代码
import random num_images = 5 random_indices = random.sample(range(len(X_test)), num_images) for idx in random_indices: image = X_test[idx] # 假设图像数组是元组的第一个元素 # print(image) # image = cv2.resize(image, (224, 224)) image1 = image.astype('float32') / 255 image1 = np.expand_dims(image1, axis=0) preds = model.predict(image1) i = np.argmax(preds[0]) icam = GradCAM(model, i, 'conv5_block3_out') heatmap = icam.compute_heatmap(image1) heatmap = cv2.resize(heatmap, (224, 224)) (heatmap, output) = icam.overlay_heatmap(heatmap, image, alpha=0.5) fig, ax = plt.subplots(1, 3) ax[0].imshow(heatmap) ax[1].imshow(image) ax[2].imshow(output)
输出异常情况

问题原因与解决方案
核心问题:梯度计算使用不可微操作
在compute_heatmap方法中,loss的计算依赖tf.argmax(predictions[0]):
loss = predictions[:, tf.argmax(predictions[0])]
tf.argmax是不可微操作,在TensorFlow的梯度跟踪上下文(tf.GradientTape)中使用时,会导致梯度计算逻辑异常,最终所有图像生成完全相同的热图。此外,这段代码完全未使用初始化时传入的self.classIdx,违背了Grad CAM针对指定类别计算激活图的设计逻辑。
修复方案
修改compute_heatmap中的loss计算代码,直接使用初始化时传入的self.classIdx:
# 替换原loss计算行 loss = predictions[:, self.classIdx]
额外优化点
- 图像类型匹配:
cv2.addWeighted要求输入图像类型一致,若X_test中的图像为float32类型,需转换为uint8:
image = X_test[idx].astype("uint8")
- 移除冗余resize:
compute_heatmap中已将热图resize至输入图像尺寸,可视化代码中无需再次执行heatmap = cv2.resize(heatmap, (224, 224)),避免尺寸不匹配问题。
修复后的compute_heatmap关键片段:
with tf.GradientTape() as tape: inputs = tf.cast(image, tf.float32) (convOutputs, predictions) = gradModel(inputs) # 使用传入的类别索引计算loss loss = predictions[:, self.classIdx]
修复后的可视化代码片段:
for idx in random_indices: # 确保图像为uint8类型 image = X_test[idx].astype("uint8") image1 = image.astype('float32') / 255 image1 = np.expand_dims(image1, axis=0) preds = model.predict(image1) i = np.argmax(preds[0]) icam = GradCAM(model, i, 'conv5_block3_out') heatmap = icam.compute_heatmap(image1) # 移除冗余resize操作 (heatmap, output) = icam.overlay_heatmap(heatmap, image, alpha=0.5) fig, ax = plt.subplots(1, 3) ax[0].imshow(heatmap) ax[1].imshow(image) ax[2].imshow(output)
内容的提问来源于stack exchange,提问作者Rezuana Haque
相关产品推荐
相关产品推荐

