Keras模型输入形状不兼容错误修复:显著性图生成问题
问题与修复方案
问题描述
作为机器学习与深度学习新手,在开展猫狗测试图像的显著性图生成项目时,已完成图像加载预处理与模型构建,但运行代码时出现输入形状不匹配错误:
Input 0 of layer "sequential" is incompatible with the layer: expected shape=(None, 300, 300, 3), found shape=(300, 300, 3),无法生成显著性图。
错误分析与修复步骤
核心错误:模型输入缺少批量维度
Keras模型默认要求输入包含批量维度(即形状为(批量大小, 高度, 宽度, 通道数),None代表可变批量大小),但代码中直接将单张图像(形状(300,300,3))传入模型,导致形状不匹配。
除此之外,代码还有以下几个小问题需要修复:
- One-hot编码逻辑错误:当前生成的
expected_output不符合单样本的标签编码要求 - 损失函数使用方式错误:CategoricalCrossentropy的调用方式不符合TensorFlow规范
- 数据类型拼写错误:
tf.unit8应为tf.uint8 - 文件名拼接错误:使用处理后的图像数组拼接文件名,导致保存失败
具体修复点
- 将
GradientTape中传给模型的输入从image改为已添加批量维度的tensor_image - 基于传入的
label正确生成one-hot编码的目标输出 - 改用函数式的
categorical_crossentropy计算损失,或正确实例化损失类 - 修正
tf.unit8为tf.uint8 - 修改函数参数名,避免覆盖原始文件名,正确拼接保存路径
修复后的完整代码
import cv2 import tensorflow as tf import matplotlib.pyplot as plt def do_salience(image_path, model, label, prefix): ''' Generates the saliency map of a given image. Args: image_path (str) -- path to the picture that the model will classify model (keras Model) -- your cats and dogs classifier label (int) -- ground truth label of the image prefix (string) -- prefix to add to the filename of the saliency map ''' # Read the image and convert channel order from BGR to RGB image = cv2.imread(image_path) image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB) # Resize the image to 300 x 300 and normalize pixel values to the range [0, 1] image = cv2.resize(image,(300,300))/255.0 # Add an additional dimension (for the batch), and save this in a new variable tensor_image = tf.expand_dims(image, axis=0) # Declare the number of classes num_classes= 2 # Define the expected output array by one-hot encoding the label expected_output = tf.one_hot([label], num_classes) # Within the GradientTape block: with tf.GradientTape() as tape: inputs = tf.cast(tensor_image, tf.float32) tape.watch(inputs) predictions = model(inputs) # 使用函数式方式计算分类交叉熵损失 loss = tf.keras.losses.categorical_crossentropy(expected_output, predictions) print(f"模型预测结果: {predictions.numpy()}") # get the gradients of the loss with respect to the model's input image gradients = tape.gradients(loss, inputs) # generate the grayscale tensor grayscale_tensor = tf.reduce_sum(tf.abs(gradients), axis=-1) # normalize the pixel values to be in the range [0, 255] min_val = tf.reduce_min(grayscale_tensor) max_val = tf.reduce_max(grayscale_tensor) normalized_tensor = tf.cast(255 * (grayscale_tensor - min_val) / (max_val - min_val), tf.uint8) # Remove dimensions that are size 1 normalized_tensor = tf.squeeze(normalized_tensor) # plot the normalized tensor plt.figure(figsize=(8, 8)) plt.axis('off') plt.imshow(normalized_tensor, cmap='gray') plt.show() # superimpose the saliency map with the original image gradient_color = cv2.applyColorMap(normalized_tensor.numpy(), cv2.COLORMAP_HOT) gradient_color = gradient_color / 255.0 super_imposed = cv2.addWeighted(image, 0.5, gradient_color, 0.5, 0.0) plt.figure(figsize=(8, 8)) plt.imshow(super_imposed) plt.axis('off') plt.show() # save the normalized tensor image to a file salient_image_name = f"{prefix}_{image_path}" normalized_tensor = tf.expand_dims(normalized_tensor, -1) normalized_tensor = tf.io.encode_jpeg(normalized_tensor, quality=100, format='grayscale') tf.io.write_file(salient_image_name, normalized_tensor) # load initial weights model.load_weights('0_epochs.h5') # generate the saliency maps for the 5 test images do_salience('cat1.jpg', model, 0, "salient") do_salience('cat2.jpg', model, 0, "salient") do_salience('catanddog.jpg', model, 0, "salient") do_salience('dog1.jpg', model, 1, "salient") do_salience('dog2.jpg', model, 1, "salient")
内容的提问来源于stack exchange,提问作者Mozhgan Zahraee
相关产品推荐
相关产品推荐

