TensorFlow训练CNN自动编码器时输入与标签形状不匹配问题求助
修复CNN自动编码器的形状不匹配问题
问题描述
在使用TensorFlow构建CNN自动编码器时,输入数据为shape=(419,128,128)的numpy数组sp_pics,运行autoencoder.fit()时出现以下错误:
ValueError:
logitsandlabelsmust have the same shape, received ((1, 124, 124, 1) vs (1, 128, 128, 1)).
原代码如下:
input_img = Input(shape=(128, 128, 1)) # Encoder x = Conv2D(16, (3, 3), activation='relu', padding='same')(input_img) x = MaxPooling2D((2, 2), padding='same')(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) x = MaxPooling2D((2, 2), padding='same')(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) encoded = MaxPooling2D((2, 2), padding='same')(x) # Decoder x = Conv2D(8, (3, 3), activation='relu', padding='same')(encoded) x = UpSampling2D((2, 2))(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) x = UpSampling2D((2, 2))(x) x = Conv2D(16, (3, 3), activation='relu')(x) x = UpSampling2D((2, 2))(x) decoded = Conv2D(1, (3, 3), activation='sigmoid', padding='same')(x) # Define Auto Encoder Model autoencoder = Model(input_img, decoded) autoencoder.compile(optimizer='adam', loss='binary_crossentropy') # Train the Autoencoder data_input = np.expand_dims(sp_pics, axis=-1) autoencoder.fit(data_input, data_input, epochs=50, batch_size=32) encoder = Model(input_img, encoded) encoded_imgs = encoder.predict(sp_pics[0,:,:]) print(encoded_imgs.shape[1])
问题原因
错误中的124来自解码器的尺寸计算偏差:
- 编码器经过3次带
padding='same'的MaxPooling2D,输入尺寸从128依次变为64→32→16,最终encoded的shape为(16,16,8) - 解码器前两次UpSampling2D后,尺寸从16恢复为32→64
- 此时经过未设置
padding='same'的Conv2D(3,3),尺寸会被压缩为64 - 3 + 1 = 62 - 最后一次UpSampling2D后,尺寸变为
62×2=124,导致输出shape为(124,124,1),与输入的(128,128,1)不匹配,触发报错。
修复方案
给解码器中缺失padding='same'的Conv2D层添加该参数,确保每一步卷积后尺寸保持不变,最终输出与输入形状一致。同时修正编码器预测时的维度错误:
import numpy as np from tensorflow.keras.layers import Input, Conv2D, MaxPooling2D, UpSampling2D from tensorflow.keras.models import Model input_img = Input(shape=(128, 128, 1)) # Encoder x = Conv2D(16, (3, 3), activation='relu', padding='same')(input_img) x = MaxPooling2D((2, 2), padding='same')(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) x = MaxPooling2D((2, 2), padding='same')(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) encoded = MaxPooling2D((2, 2), padding='same')(x) # Decoder - 修复Conv2D的padding参数 x = Conv2D(8, (3, 3), activation='relu', padding='same')(encoded) x = UpSampling2D((2, 2))(x) x = Conv2D(8, (3, 3), activation='relu', padding='same')(x) x = UpSampling2D((2, 2))(x) x = Conv2D(16, (3, 3), activation='relu', padding='same')(x) # 添加padding='same' x = UpSampling2D((2, 2))(x) decoded = Conv2D(1, (3, 3), activation='sigmoid', padding='same')(x) # Define Auto Encoder Model autoencoder = Model(input_img, decoded) autoencoder.compile(optimizer='adam', loss='binary_crossentropy') # Train the Autoencoder data_input = np.expand_dims(sp_pics, axis=-1) autoencoder.fit(data_input, data_input, epochs=50, batch_size=32) encoder = Model(input_img, encoded) # 修正输入维度:扩展为模型要求的4维格式(批量, 高, 宽, 通道) encoded_imgs = encoder.predict(np.expand_dims(sp_pics[0,:,:], axis=(0, -1))) print(encoded_imgs.shape[1]) # 输出应为16
内容的提问来源于stack exchange,提问作者Manfred
相关产品推荐
相关产品推荐

