针对任意CNN架构的回归激活映射(RAM)实现及代码问询
Hey there! Let's tackle your two questions step by step—first, a practical code example for Regression Activation Mapping (RAM), then adjusting those lines when dealing with scalar predictions.
1. RAM Code Example (Keras/TensorFlow)
Since RAM is built for regression tasks (replacing classification-focused CAM), we'll adapt the Grad-CAM framework to work with continuous output values. This example aligns with the diabetic retinopathy detection use case you mentioned, following the core logic from the referenced paper:
import numpy as np import tensorflow as tf from tensorflow.keras.models import Model from tensorflow.keras.preprocessing.image import load_img, img_to_array import cv2 import matplotlib.pyplot as plt def compute_ram(model, img_array, last_conv_layer_name): # Build a model that links input to last conv layer activations + regression output grad_model = Model( inputs=model.inputs, outputs=[model.get_layer(last_conv_layer_name).output, model.output] ) # Calculate gradients of the scalar regression output w.r.t. conv layer activations with tf.GradientTape() as tape: conv_outputs, pred = grad_model(img_array) # For regression, the loss is just the scalar prediction itself loss = pred[0] # Compute gradients of the loss against the conv layer outputs grads = tape.gradient(loss, conv_outputs) # Average pool gradients across spatial dimensions pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2)) # Weight conv layer activations by pooled gradients to get RAM heatmap conv_outputs = conv_outputs[0] ram_heatmap = tf.reduce_sum(conv_outputs * pooled_grads, axis=-1) # Normalize heatmap to [0,1] for visualization ram_heatmap = np.maximum(ram_heatmap, 0) ram_heatmap /= np.max(ram_heatmap) return ram_heatmap # Example Usage # Assume your diabetic retinopathy model is loaded (e.g., from the repo you referenced) # model = tf.keras.models.load_model('path/to/your/model.h5') # last_conv_layer = 'your_last_convolutional_layer_name' # Load and preprocess a retina image img = load_img('retina_sample.jpg', target_size=(224, 224)) img_array = img_to_array(img) img_array = np.expand_dims(img_array, axis=0) # Add any preprocessing your model requires (e.g., scaling) # img_array = preprocess_input(img_array) # Generate RAM heatmap ram_map = compute_ram(model, img_array, last_conv_layer) # Optional: Overlay heatmap on original image original_img = cv2.imread('retina_sample.jpg') original_img = cv2.resize(original_img, (ram_map.shape[1], ram_map.shape[0])) heatmap = cv2.applyColorMap(np.uint8(255 * ram_map), cv2.COLORMAP_JET) superimposed_img = heatmap * 0.4 + original_img cv2.imwrite('ram_overlay.jpg', superimposed_img)
Key RAM Details:
- Unlike CAM/Grad-CAM, RAM doesn't use class indices—we directly use the scalar regression output as our loss signal.
- The heatmap highlights regions in the image that most influence the continuous prediction (e.g., retinopathy severity score).
- This implementation matches the paper's goal of adapting activation mapping for non-classification tasks.
2. Adjusting Code for Scalar Predictions
Your original code lines are designed for classification tasks where preds is a vector of class probabilities. When dealing with a scalar regression output, you can skip the class index selection entirely:
Original Classification Code:
class_idx = np.argmax(preds[0]) class_output = model.output[:, class_idx]
Modified Code for Scalar pred:
# No class index needed for scalar regression output regression_output = model.output # Directly use the scalar output tensor # If you need to access the prediction value (for loss calculation), use: # loss = pred[0]
For RAM/Grad-CAM workflows with scalar outputs, you don't need to pick a class—just compute gradients against the model's raw scalar output to find which image regions drive the prediction.
内容的提问来源于stack exchange,提问作者Endre Moen

