如何向TensorFlow TimeDistributed层输入可变数量的图像?
我正在用TensorFlow构建以EfficientNet为Backbone的CNN进行图像分类。原始图像尺寸大且包含大量空白区域,因此被分割成tiles(图像块),比如第一张图像有30个256×256的tiles。但并非所有图像的tiles数量都固定为30,有的多有的少,且所有tiles都很重要。我打算用LSTM(配合TimeDistributed层)处理这种可变长度输入。
TensorFlow Dataset输出的数据形状为(None, 256, 256, 3),其中None代表tiles数量可变,后三维对应每个tiles的(256,256,3)形状。
由于EfficientNet的input_shape仅支持3个输入通道,我将其设置为(256, 256, 3),代码如下:
B0 = tf.keras.applications.efficientnet_v2.EfficientNetV2B0( weights='imagenet', include_top=False, pooling='avg', input_shape=(256, 256, 3), )
参考固定tiles数量的示例(每张图像固定10个tiles),我尝试构建支持可变tiles数量的模型:
B0 = tf.keras.applications.efficientnet_v2.EfficientNetV2B0( weights='imagenet', include_top=False, pooling='avg', input_shape=(TILE_SIZE, TILE_SIZE, 3), ) B0 = Model(inputs=B0.inputs, outputs=B0.layers[-2].output) model = Sequential() model.add(TimeDistributed(B0, input_shape=(None, TILE_SIZE, TILE_SIZE, 3))) model.add(GlobalMaxPooling3D()) model.add(Dense(6, activation='softmax')) model.compile( loss='categorical_crossentropy', optimizer=Adam(), metrics=['categorical_accuracy', tfa.metrics.CohenKappa(num_classes=6, sparse_labels=False, weightage="quadratic")] )
运行后出现错误:
WARNING:tensorflow:Model was constructed with shape (None, None, 256, 256, 3) for input KerasTensor(type_spec=TensorSpec(shape=(None, None, 256, 256, 3), dtype=tf.float32, name='time_distributed_input'), name='time_distributed_input', description="created by layer 'time_distributed_input'"), but it was called on an input with incompatible shape (None, None, None, None). Traceback (most recent call last): File "C:\Users\tom50\OneDrive - Flinders\Honours\TimeDistributed\Main.py", line 59, in <module> history = model.fit( File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\utils\traceback_utils.py", line 70, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\tom50\AppData\Local\Temp\__autograph_generated_filekcf5gc2z.py", line 15, in tf__train_function retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope) ValueError: in user code: File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\engine\training.py", line 1160, in train_function * return step_function(self, iterator) File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\engine\training.py", line 1146, in step_function ** outputs = model.distribute_strategy.run(run_step, args=(data,)) File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\engine\training.py", line 1135, in run_step ** outputs = model.train_step(data) File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\engine\training.py", line 993, in train_step y_pred = self(x, training=True) File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\utils\traceback_utils.py", line 70, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\tom50\AppData\Roaming\Python\Python39\site-packages\keras\engine\input_spec.py", line 232, in assert_input_compatibility raise ValueError( ValueError: Exception encountered when calling layer "sequential" " f"(type Sequential). Input 0 of layer "time_distributed" is incompatible with the layer: expected ndim=5, found ndim=4. Full shape received: (None, None, None, None) Call arguments received by layer "sequential" " f"(type Sequential): • inputs=tf.Tensor(shape=(None, None, None, None), dtype=float32) • training=True • mask=None
输入维度不匹配:TimeDistributed层期望接收5维输入(
(batch_size, seq_len, height, width, channels)),但实际输入是4维((None, None, None, None))。问题出在TensorFlow Dataset的输出格式——当前Dataset输出的每个样本是4维((tiles_num, 256,256,3)),但batch后整体形状应为5维((batch_size, tiles_num,256,256,3)),而实际加载时没有正确构建这种5维结构。池化层使用错误:GlobalMaxPooling3D需要5维输入(
(batch, dim1, dim2, dim3, channels)),但TimeDistributed包裹EfficientNet后的输出是3维((batch_size, seq_len, feature_dim)),两者维度不兼容。
1. 修正TensorFlow Dataset的输出形状
确保Dataset输出的每个样本是可变长度的tiles序列,batch后形成5维张量。例如,使用tf.data.Dataset.from_generator加载数据时,将每个图像的所有tiles打包成一个形状为(None,256,256,3)的张量,再进行batch操作:
def load_data_generator(): # 遍历每个图像的tiles数据 for img_tiles, label in your_data_source: # img_tiles形状为(tiles_num,256,256,3) yield tf.convert_to_tensor(img_tiles, dtype=tf.float32), tf.convert_to_tensor(label, dtype=tf.float32) dataset = tf.data.Dataset.from_generator( load_data_generator, output_signature=( tf.TensorSpec(shape=(None,256,256,3), dtype=tf.float32), tf.TensorSpec(shape=(6,), dtype=tf.float32) # 对应分类标签 ) ) dataset = dataset.batch(batch_size=8) # batch后形状为(8, None,256,256,3)
2. 调整模型结构,替换池化层
将GlobalMaxPooling3D替换为适合处理序列特征的GlobalMaxPooling1D,或者添加LSTM层提取序列信息:
# 方案一:使用GlobalMaxPooling1D B0 = tf.keras.applications.efficientnet_v2.EfficientNetV2B0( weights='imagenet', include_top=False, pooling='avg', input_shape=(256, 256, 3), ) B0 = tf.keras.Model(inputs=B0.inputs, outputs=B0.layers[-2].output) model = tf.keras.Sequential() model.add(tf.keras.layers.TimeDistributed(B0, input_shape=(None, 256, 256, 3))) model.add(tf.keras.layers.GlobalMaxPooling1D()) # 处理序列维度 model.add(tf.keras.layers.Dense(6, activation='softmax')) # 方案二:使用LSTM提取序列特征 model = tf.keras.Sequential() model.add(tf.keras.layers.TimeDistributed(B0, input_shape=(None, 256, 256, 3))) model.add(tf.keras.layers.LSTM(128, return_sequences=False)) # 提取序列特征 model.add(tf.keras.layers.Dense(6, activation='softmax'))
3. 验证输入形状
在训练前可以打印Dataset的元素形状,确认输入符合模型要求:
for x, y in dataset.take(1): print("Input shape:", x.shape) # 应输出类似(8, None,256,256,3)的5维形状
内容的提问来源于stack exchange,提问作者Tom Lin

