You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ResNet50时Grad CAM输出所有图像热图一致问题求助

Grad CAM热图异常:所有图像输出一致的问题排查与解决

我使用Grad CAM分析ResNet50模型对测试图像的预测关键区域,但出现异常:所有测试图像的原始图各不相同,但对应的热图和Grad CAM叠加结果完全一致。

Grad CAM实现代码

from tensorflow.keras.models import Model
import tensorflow as tf
import numpy as np
import cv2

class GradCAM:
    def __init__(self, model, classIdx, layerName=None):
        # 保存模型、用于计算类激活图的类别索引,以及可视化目标层
        self.model = model
        self.classIdx = classIdx
        self.layerName = layerName
        # 如果未指定层名,自动查找目标输出层
        if self.layerName is None:
            self.layerName = self.find_target_layer()

    def find_target_layer(self):
        # 反向遍历网络层,寻找最后一个4D输出的卷积层
        for layer in reversed(self.model.layers):
            if len(layer.output_shape) == 4:
                return layer.name
        # 未找到4D层则无法应用GradCAM
        raise ValueError("Could not find 4D layer. Cannot apply GradCAM.")


    def compute_heatmap(self, image, eps=1e-8):
        # 构建梯度模型:输入为原模型输入,输出为目标卷积层输出和模型最终预测
        gradModel = Model(
            inputs=[self.model.inputs],
            outputs=[self.model.get_layer(self.layerName).output, self.model.output])

        # 记录自动微分操作
        with tf.GradientTape() as tape:
            inputs = tf.cast(image, tf.float32)
            (convOutputs, predictions) = gradModel(inputs)
            
            loss = predictions[:, tf.argmax(predictions[0])]
    
        # 计算梯度
        grads = tape.gradient(loss, convOutputs)

        # 计算引导梯度
        castConvOutputs = tf.cast(convOutputs > 0, "float32")
        castGrads = tf.cast(grads > 0, "float32")
        guidedGrads = castConvOutputs * castGrads * grads
        # 去除batch维度
        convOutputs = convOutputs[0]
        guidedGrads = guidedGrads[0]

        # 计算梯度均值作为权重,加权求和得到类激活图
        weights = tf.reduce_mean(guidedGrads, axis=(0, 1))
        cam = tf.reduce_sum(tf.multiply(weights, convOutputs), axis=-1)

        # 将类激活图resize到输入图像尺寸
        (w, h) = (image.shape[2], image.shape[1])
        heatmap = cv2.resize(cam.numpy(), (w, h))
        # 归一化热图到[0,255]区间
        numer = heatmap - np.min(heatmap)
        denom = (heatmap.max() - heatmap.min()) + eps
        heatmap = numer / denom
        heatmap = (heatmap * 255).astype("uint8")
        return heatmap

    def overlay_heatmap(self, heatmap, image, alpha=0.5,
                        colormap=cv2.COLORMAP_VIRIDIS):
        # 为热图添加颜色映射并与原图叠加
        heatmap = cv2.applyColorMap(heatmap, colormap)
        output = cv2.addWeighted(image, alpha, heatmap, 1 - alpha, 0)
        return (heatmap, output)

热图可视化代码

import random

num_images = 5
random_indices = random.sample(range(len(X_test)), num_images)

for idx in random_indices:
    image = X_test[idx] # 假设图像数组是元组的第一个元素
    # print(image)
    # image = cv2.resize(image, (224, 224))
    image1 = image.astype('float32') / 255
    image1 = np.expand_dims(image1, axis=0)
    preds = model.predict(image1) 
    i = np.argmax(preds[0])
    icam = GradCAM(model, i, 'conv5_block3_out') 
    heatmap = icam.compute_heatmap(image1)
    heatmap = cv2.resize(heatmap, (224, 224))
    (heatmap, output) = icam.overlay_heatmap(heatmap, image, alpha=0.5)
    fig, ax = plt.subplots(1, 3)
    ax[0].imshow(heatmap)
    ax[1].imshow(image)
    ax[2].imshow(output)

输出异常情况

输出异常的Grad CAM结果

问题原因与解决方案

核心问题:梯度计算使用不可微操作

在compute_heatmap方法中,loss的计算依赖tf.argmax(predictions[0]):

loss = predictions[:, tf.argmax(predictions[0])]

tf.argmax是不可微操作,在TensorFlow的梯度跟踪上下文(tf.GradientTape)中使用时,会导致梯度计算逻辑异常,最终所有图像生成完全相同的热图。此外,这段代码完全未使用初始化时传入的self.classIdx,违背了Grad CAM针对指定类别计算激活图的设计逻辑。

修复方案

修改compute_heatmap中的loss计算代码,直接使用初始化时传入的self.classIdx:

# 替换原loss计算行
loss = predictions[:, self.classIdx]

额外优化点

  1. 图像类型匹配:cv2.addWeighted要求输入图像类型一致,若X_test中的图像为float32类型,需转换为uint8:
image = X_test[idx].astype("uint8")
  1. 移除冗余resize:compute_heatmap中已将热图resize至输入图像尺寸,可视化代码中无需再次执行heatmap = cv2.resize(heatmap, (224, 224)),避免尺寸不匹配问题。

修复后的compute_heatmap关键片段:

with tf.GradientTape() as tape:
    inputs = tf.cast(image, tf.float32)
    (convOutputs, predictions) = gradModel(inputs)
    # 使用传入的类别索引计算loss
    loss = predictions[:, self.classIdx]

修复后的可视化代码片段:

for idx in random_indices:
    # 确保图像为uint8类型
    image = X_test[idx].astype("uint8")
    image1 = image.astype('float32') / 255
    image1 = np.expand_dims(image1, axis=0)
    preds = model.predict(image1) 
    i = np.argmax(preds[0])
    icam = GradCAM(model, i, 'conv5_block3_out') 
    heatmap = icam.compute_heatmap(image1)
    # 移除冗余resize操作
    (heatmap, output) = icam.overlay_heatmap(heatmap, image, alpha=0.5)
    fig, ax = plt.subplots(1, 3)
    ax[0].imshow(heatmap)
    ax[1].imshow(image)
    ax[2].imshow(output)

内容的提问来源于stack exchange,提问作者Rezuana Haque

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 20:15:55