基于Keras的CNN类别可视化与Google Dream风格图像生成技术问询
Hey there! Since you already have a fine-tuned InceptionV3 model (inceptionv3-ft.model) ready to go, let's dive into practical, actionable solutions for both your needs:
1. CNN Class Visualization with Grad-CAM
The most effective way to visualize which regions of an image drive your model's class predictions is Grad-CAM (Gradient-weighted Class Activation Mapping). It works seamlessly with your fine-tuned model—no retraining or architecture modifications required. Here's how to implement it:
Step-by-Step Implementation
- Load your fine-tuned model:
from tensorflow.keras.models import load_model model = load_model('inceptionv3-ft.model') - Identify key layers: Check your
model.summary()output to find the last convolutional layer (for InceptionV3, this is typically something like'mixed10'or a similar inception module). - Build a Grad-CAM model: We'll create a dual-output model to capture both the last conv layer's activations and the model's final predictions:
import tensorflow as tf last_conv_layer_name = "mixed10" # Update this to match your model's layer name classifier_layer_names = [layer.name for layer in model.layers[model.layers.index(model.get_layer(last_conv_layer_name))+1:]] # Model to output last conv layer activations last_conv_layer_model = tf.keras.Model(model.inputs, model.get_layer(last_conv_layer_name).output) # Model to map conv layer outputs to final predictions classifier_input = tf.keras.Input(shape=last_conv_layer_model.output.shape[1:]) x = classifier_input for layer_name in classifier_layer_names: x = model.get_layer(layer_name)(x) classifier_model = tf.keras.Model(classifier_input, x) - Compute the Grad-CAM heatmap:
def make_gradcam_heatmap(img_array, class_index=None): with tf.GradientTape() as tape: conv_outputs = last_conv_layer_model(img_array) tape.watch(conv_outputs) preds = classifier_model(conv_outputs) if class_index is None: class_index = tf.argmax(preds[0]) class_channel = preds[:, class_index] # Calculate gradients of the target class w.r.t. conv layer outputs grads = tape.gradient(class_channel, conv_outputs) pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2)) # Weight conv layer outputs by gradient importance heatmap = conv_outputs[0] @ pooled_grads[..., tf.newaxis] heatmap = tf.squeeze(heatmap) # Normalize for visualization heatmap = tf.maximum(heatmap, 0) / tf.math.reduce_max(heatmap) return heatmap.numpy() - Overlay heatmap on the original image:
import matplotlib.pyplot as plt import numpy as np from PIL import Image # Preprocess input image (match your training preprocessing) img = Image.open("your_input_image.jpg") img = img.resize((299, 299)) # InceptionV3's default input size img_array = tf.keras.preprocessing.image.img_to_array(img) img_array = np.expand_dims(img_array, axis=0) img_array = tf.keras.applications.inception_v3.preprocess_input(img_array) # Generate heatmap heatmap = make_gradcam_heatmap(img_array) # Convert heatmap to RGB heatmap = np.uint8(255 * heatmap) jet = plt.cm.get_cmap("jet") jet_colors = jet(np.arange(256))[:, :3] jet_heatmap = jet_colors[heatmap] # Superimpose heatmap on original image jet_heatmap = tf.keras.preprocessing.image.array_to_img(jet_heatmap) jet_heatmap = jet_heatmap.resize(img.size) jet_heatmap = tf.keras.preprocessing.image.img_to_array(jet_heatmap) superimposed_img = jet_heatmap * 0.4 + img_array[0] superimposed_img = tf.keras.preprocessing.image.array_to_img(superimposed_img) # Save or display the result superimposed_img.save("gradcam_result.jpg") plt.imshow(superimposed_img) plt.show()
2. DeepDream-Style Image Generation with Gradient Ascent
To create those surreal, dream-like images, we'll use gradient ascent to maximize the activation of specific layers in your fine-tuned InceptionV3 model. Here's how to adapt this to your setup:
Core Idea
We define a loss function that maximizes the sum of activations from mid-level layers (these capture a mix of simple and complex features), then iteratively update the input image to boost this loss.
Step-by-Step Implementation
- Select target layers: Use
model.summary()to pick layers like'mixed3','mixed4', or'mixed5'—these work best for balanced DeepDream effects. - Define the loss function:
def compute_loss(input_image, target_layers): loss = tf.zeros(shape=()) for layer_name in target_layers: layer = model.get_layer(layer_name) activation = layer(input_image) loss += tf.reduce_mean(tf.square(activation)) # Maximize L2 norm of activations return loss - Gradient ascent loop:
@tf.function def gradient_ascent_step(img, target_layers, step_size): with tf.GradientTape() as tape: tape.watch(img) loss = compute_loss(img, target_layers) grads = tape.gradient(loss, img) grads = tf.math.l2_normalize(grads) # Stabilize ascent with normalized gradients img += step_size * grads return loss, img def deepdream(img, target_layers, iterations=100, step_size=0.01): img = tf.convert_to_tensor(img) img = tf.keras.applications.inception_v3.preprocess_input(img) for i in range(iterations): loss, img = gradient_ascent_step(img, target_layers, step_size) if i % 10 == 0: print(f"Iteration {i}, Loss: {loss.numpy():.4f}") # Convert back to a displayable image img = img + 1.0 img = img / 2.0 img = img * 255.0 img = tf.cast(img, tf.uint8) return img.numpy() - Run DeepDream on your base image:
# Load and preprocess your base image base_img = Image.open("your_base_image.jpg") base_img = base_img.resize((299, 299)) base_img_array = tf.keras.preprocessing.image.img_to_array(base_img) base_img_array = np.expand_dims(base_img_array, axis=0) # Choose target layers (adjust for different effects) target_layers = ["mixed3", "mixed5"] # Generate the dream image dream_img = deepdream(base_img_array, target_layers, iterations=150, step_size=0.02) # Save or display the result dream_pil = Image.fromarray(dream_img[0]) dream_pil.save("deepdream_result.jpg") plt.imshow(dream_pil) plt.show()
Pro Tips for Better Results
- Multi-scale processing: Run gradient ascent on increasingly scaled versions of the image to add finer details.
- Gaussian blur: Blur the image between iterations to reduce noise and smooth patterns.
- Layer tuning: Lower layers (like
mixed2) produce simple geometric shapes, while higher layers (likemixed10) generate more complex, object-like patterns.
内容的提问来源于stack exchange,提问作者user7830303

