加载自定义预训练TensorFlow模型时出现负维度尺寸错误
解决自定义Xception模型替换官方预训练模型后的维度错误
问题重现
使用Keras官方预训练Xception模型时,以下代码可正常运行;但替换为自行训练的模型(通过tf.keras.models.load_model('/content/drive/MyDrive/MLC')加载)后,抛出维度错误:
def create_vision_encoder( num_projection_layers, projection_dims, dropout_rate, trainable=False ): # 加载预训练模型 xception = keras.applications.Xception( include_top=False, weights="imagenet", pooling="avg" ) for layer in xception.layers: layer.trainable = trainable inputs = layers.Input(shape=(299, 299, 3), name="image_input") BATCH_SIZE = 1 NUM_BOXES = 5 IMAGE_HEIGHT = 256 IMAGE_WIDTH = 256 CHANNELS = 3 CROP_SIZE = (24, 24) boxes = tf.random.uniform(shape=(NUM_BOXES, 4)) box_indices = tf.random.uniform(shape=(NUM_BOXES,), minval=0, maxval=BATCH_SIZE, dtype=tf.int32) output = tf.image.crop_and_resize(inputs, boxes, box_indices, CROP_SIZE) xception_input = tf.keras.applications.xception.preprocess_input(output) embeddings = xception(xception_input) outputs = project_embeddings( embeddings, num_projection_layers, projection_dims, dropout_rate ) return keras.Model(inputs, outputs, name="vision_encoder")
错误信息:
ValueError: Exception encountered when calling layer "max_pooling2d_1" (type MaxPooling2D). Negative dimension size caused by subtracting 3 from 2 for '{{node sequential/max_pooling2d_1/MaxPool}} = MaxPool[T=DT_FLOAT, data_format="NHWC", explicit_paddings=[], ksize=[1, 3, 3, 1], padding="VALID", strides=[1, 2, 2, 1]](sequential/batch_normalization_1/FusedBatchNormV3)' with input shapes: [5,2,2,256]. Call arguments received: • inputs=tf.Tensor(shape=(5, 2, 2, 256), dtype=float32)
错误原因
自定义模型中的max_pooling2d_1层使用了VALID填充、3x3池化核与2步长,而经过前序卷积层处理后,输入到该池化层的特征图尺寸仅为2x2。按照VALID填充的计算逻辑:输出尺寸 = ceil((输入尺寸 - 池化核尺寸 + 1)/步长),代入后得到负数,导致维度错误。
官方预训练Xception模型未出现该问题,是因为其网络结构针对299x299输入设计,即便接收24x24输入,前序层处理后的特征图尺寸仍能满足池化层的要求;而自行训练的模型可能在训练时使用了更大的输入尺寸,未适配小尺寸输入的情况。
解决方案
1. 增大裁剪尺寸
将CROP_SIZE从(24,24)调整为更大的值(如64x64、128x128),确保经过前序卷积层后,输入到池化层的特征图尺寸≥3x3。例如:
CROP_SIZE = (64, 64)
2. 修改池化层填充方式
加载自定义模型后,将出错的MaxPooling2D层的填充方式从VALID改为SAME,自动填充边缘以避免负维度:
xception = tf.keras.models.load_model('/content/drive/MyDrive/MLC') # 定位出错的池化层 for layer in xception.layers: if layer.name == "max_pooling2d_1": layer.padding = "same" # 重新构建层的计算图 layer.build(layer.input_shape) # 重新编译模型(根据需求调整优化器、损失函数) xception.compile(optimizer="adam", loss="categorical_crossentropy")
3. 对裁剪后的图像进行上采样
在输入自定义模型前,将24x24的裁剪图放大到模型训练时的输入尺寸(如299x299):
xception_input = tf.keras.applications.xception.preprocess_input(output) # 调整到训练时的输入尺寸 xception_input = tf.keras.layers.Resizing(299, 299)(xception_input) embeddings = xception(xception_input)
4. 对齐训练与推理的输入尺寸
检查自定义模型训练时使用的输入尺寸,确保推理时裁剪后的图像尺寸与训练尺寸一致,从根源避免维度不匹配问题。
内容的提问来源于stack exchange,提问作者albert
相关产品推荐
相关产品推荐

