You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow训练CNN自动编码器时输入与标签形状不匹配问题求助

修复CNN自动编码器的形状不匹配问题

问题描述

在使用TensorFlow构建CNN自动编码器时,输入数据为shape=(419,128,128)的numpy数组sp_pics,运行autoencoder.fit()时出现以下错误:

ValueError: logits and labels must have the same shape, received ((1, 124, 124, 1) vs (1, 128, 128, 1)).

原代码如下:

input_img = Input(shape=(128, 128, 1))

# Encoder
x = Conv2D(16, (3, 3), activation='relu', padding='same')(input_img)
x = MaxPooling2D((2, 2), padding='same')(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
x = MaxPooling2D((2, 2), padding='same')(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
encoded = MaxPooling2D((2, 2), padding='same')(x)

# Decoder
x = Conv2D(8, (3, 3), activation='relu', padding='same')(encoded)
x = UpSampling2D((2, 2))(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
x = UpSampling2D((2, 2))(x)
x = Conv2D(16, (3, 3), activation='relu')(x)
x = UpSampling2D((2, 2))(x)
decoded = Conv2D(1, (3, 3), activation='sigmoid', padding='same')(x)

# Define Auto Encoder Model
autoencoder = Model(input_img, decoded)

autoencoder.compile(optimizer='adam', loss='binary_crossentropy')

# Train the Autoencoder
data_input = np.expand_dims(sp_pics, axis=-1)
autoencoder.fit(data_input, data_input, epochs=50, batch_size=32)

encoder = Model(input_img, encoded)
encoded_imgs = encoder.predict(sp_pics[0,:,:])

print(encoded_imgs.shape[1])

问题原因

错误中的124来自解码器的尺寸计算偏差:

  1. 编码器经过3次带padding='same'的MaxPooling2D,输入尺寸从128依次变为64→32→16,最终encoded的shape为(16,16,8)
  2. 解码器前两次UpSampling2D后,尺寸从16恢复为32→64
  3. 此时经过未设置padding='same'的Conv2D(3,3),尺寸会被压缩为64 - 3 + 1 = 62
  4. 最后一次UpSampling2D后,尺寸变为62×2=124,导致输出shape为(124,124,1),与输入的(128,128,1)不匹配,触发报错。

修复方案

给解码器中缺失padding='same'的Conv2D层添加该参数,确保每一步卷积后尺寸保持不变,最终输出与输入形状一致。同时修正编码器预测时的维度错误:

import numpy as np
from tensorflow.keras.layers import Input, Conv2D, MaxPooling2D, UpSampling2D
from tensorflow.keras.models import Model

input_img = Input(shape=(128, 128, 1))

# Encoder
x = Conv2D(16, (3, 3), activation='relu', padding='same')(input_img)
x = MaxPooling2D((2, 2), padding='same')(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
x = MaxPooling2D((2, 2), padding='same')(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
encoded = MaxPooling2D((2, 2), padding='same')(x)

# Decoder - 修复Conv2D的padding参数
x = Conv2D(8, (3, 3), activation='relu', padding='same')(encoded)
x = UpSampling2D((2, 2))(x)
x = Conv2D(8, (3, 3), activation='relu', padding='same')(x)
x = UpSampling2D((2, 2))(x)
x = Conv2D(16, (3, 3), activation='relu', padding='same')(x)  # 添加padding='same'
x = UpSampling2D((2, 2))(x)
decoded = Conv2D(1, (3, 3), activation='sigmoid', padding='same')(x)

# Define Auto Encoder Model
autoencoder = Model(input_img, decoded)

autoencoder.compile(optimizer='adam', loss='binary_crossentropy')

# Train the Autoencoder
data_input = np.expand_dims(sp_pics, axis=-1)
autoencoder.fit(data_input, data_input, epochs=50, batch_size=32)

encoder = Model(input_img, encoded)
# 修正输入维度:扩展为模型要求的4维格式(批量, 高, 宽, 通道)
encoded_imgs = encoder.predict(np.expand_dims(sp_pics[0,:,:], axis=(0, -1)))

print(encoded_imgs.shape[1])  # 输出应为16

内容的提问来源于stack exchange,提问作者Manfred

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:40:22