You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras模型输入形状不兼容错误修复:显著性图生成问题

问题与修复方案

问题描述

作为机器学习与深度学习新手,在开展猫狗测试图像的显著性图生成项目时,已完成图像加载预处理与模型构建,但运行代码时出现输入形状不匹配错误:
Input 0 of layer "sequential" is incompatible with the layer: expected shape=(None, 300, 300, 3), found shape=(300, 300, 3),无法生成显著性图。

错误分析与修复步骤

核心错误:模型输入缺少批量维度

Keras模型默认要求输入包含批量维度(即形状为(批量大小, 高度, 宽度, 通道数),None代表可变批量大小),但代码中直接将单张图像(形状(300,300,3))传入模型,导致形状不匹配。

除此之外,代码还有以下几个小问题需要修复:

  • One-hot编码逻辑错误:当前生成的expected_output不符合单样本的标签编码要求
  • 损失函数使用方式错误:CategoricalCrossentropy的调用方式不符合TensorFlow规范
  • 数据类型拼写错误:tf.unit8应为tf.uint8
  • 文件名拼接错误:使用处理后的图像数组拼接文件名,导致保存失败

具体修复点

  1. 将GradientTape中传给模型的输入从image改为已添加批量维度的tensor_image
  2. 基于传入的label正确生成one-hot编码的目标输出
  3. 改用函数式的categorical_crossentropy计算损失,或正确实例化损失类
  4. 修正tf.unit8为tf.uint8
  5. 修改函数参数名,避免覆盖原始文件名,正确拼接保存路径

修复后的完整代码

import cv2
import tensorflow as tf
import matplotlib.pyplot as plt

def do_salience(image_path, model, label, prefix):
  '''
  Generates the saliency map of a given image.

  Args:
    image_path (str) -- path to the picture that the model will classify
    model (keras Model) -- your cats and dogs classifier
    label (int) -- ground truth label of the image
    prefix (string) -- prefix to add to the filename of the saliency map
  '''

  # Read the image and convert channel order from BGR to RGB
  image = cv2.imread(image_path)
  image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB) 
  # Resize the image to 300 x 300 and normalize pixel values to the range [0, 1]
  image = cv2.resize(image,(300,300))/255.0

  # Add an additional dimension (for the batch), and save this in a new variable
  tensor_image = tf.expand_dims(image, axis=0)

  # Declare the number of classes
  num_classes= 2

  # Define the expected output array by one-hot encoding the label
  expected_output = tf.one_hot([label], num_classes)

  # Within the GradientTape block:
  with tf.GradientTape() as tape:
    inputs = tf.cast(tensor_image, tf.float32)
    tape.watch(inputs)
    predictions = model(inputs)
    # 使用函数式方式计算分类交叉熵损失
    loss = tf.keras.losses.categorical_crossentropy(expected_output, predictions)
    print(f"模型预测结果: {predictions.numpy()}")
  
  # get the gradients of the loss with respect to the model's input image
  gradients = tape.gradients(loss, inputs)
  
  # generate the grayscale tensor
  grayscale_tensor = tf.reduce_sum(tf.abs(gradients), axis=-1)

  # normalize the pixel values to be in the range [0, 255]
  min_val = tf.reduce_min(grayscale_tensor)
  max_val = tf.reduce_max(grayscale_tensor)
  normalized_tensor = tf.cast(255 * (grayscale_tensor - min_val) / (max_val - min_val), tf.uint8)
    
  # Remove dimensions that are size 1
  normalized_tensor = tf.squeeze(normalized_tensor)
    
  # plot the normalized tensor
  plt.figure(figsize=(8, 8))
  plt.axis('off')
  plt.imshow(normalized_tensor, cmap='gray')
  plt.show()

  # superimpose the saliency map with the original image
  gradient_color = cv2.applyColorMap(normalized_tensor.numpy(), cv2.COLORMAP_HOT)
  gradient_color = gradient_color / 255.0
  super_imposed = cv2.addWeighted(image, 0.5, gradient_color, 0.5, 0.0)
  plt.figure(figsize=(8, 8))
  plt.imshow(super_imposed)
  plt.axis('off')
  plt.show()
  
  # save the normalized tensor image to a file
  salient_image_name = f"{prefix}_{image_path}"
  normalized_tensor = tf.expand_dims(normalized_tensor, -1)
  normalized_tensor = tf.io.encode_jpeg(normalized_tensor, quality=100, format='grayscale')
  tf.io.write_file(salient_image_name, normalized_tensor)

# load initial weights
model.load_weights('0_epochs.h5')

# generate the saliency maps for the 5 test images
do_salience('cat1.jpg', model, 0, "salient")
do_salience('cat2.jpg', model, 0, "salient")
do_salience('catanddog.jpg', model, 0, "salient")
do_salience('dog1.jpg', model, 1, "salient")
do_salience('dog2.jpg', model, 1, "salient")

内容的提问来源于stack exchange,提问作者Mozhgan Zahraee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:47:52