使用tf.data.Dataset时model.fit无法识别标签的问题
问题描述
通过生成器创建TensorFlow数据集,为自编码器添加固定标量0作为标签,训练时触发ValueError,报错提示切片索引越界。
代码与报错详情
数据集创建与标签添加
tfds = tf.data.Dataset.from_generator(streamFromFile, output_signature=tf.TensorSpec(shape=(240,120), dtype=tf.uint16)) def prepareLabels(x): return x, 0 tfds = tfds.map(prepareLabels) # 输出验证 for x,y in tfds.take(1): print("features", x.shape) print("label", y.shape) # 输出: # features (240, 120) # label ()
训练代码
model.fit( tfds.take(10).batch(4), epochs=3 )
报错信息
ValueError: in user code: File "/home/user/.local/lib/python3.11/site-packages/keras/src/engine/training.py", line 1401, in train_function * return step_function(self, iterator) File "/home/user/.local/lib/python3.11/site-packages/keras/src/engine/training.py", line 1384, in step_function ** outputs = model.distribute_strategy.run(run_step, args=(data,)) ... is_last_dim_1 = tf.equal(1, tf.shape(y_pred)[-1]) ValueError: slice index -1 of dimension 0 out of bounds. for '{{node mean_absolute_error/strided_slice}} = StridedSlice[Index=DT_INT32, T=DT_INT32, begin_mask=0, ellipsis_mask=0, end_mask=0, new_axis_mask=0, shrink_axis_mask=1](mean_absolute_error/Shape, mean_absolute_error/strided_slice/stack, mean_absolute_error/strided_slice/stack_1, mean_absolute_error/strided_slice/stack_2)' with input shapes: [0], [1], [1], [1] and with computed input tensors: input[1] = <-1>, input[2] = <0>, input[3] = <1>.
模型结构摘要
__________________________________________________________________________________________________ Layer (type) Output Shape Param # Connected to ================================================================================================== input_3 (InputLayer) [(None, 240, 120, 1)] 0 [] ... encoded (Dense) (None, 200) 40200 ['reshape_4[0][0]'] ... reconstruction (Conv2DTran (None, 240, 120, 1) 37 ['up_sampling2d_11[0][0]'] spose) score (ScoreLayer) () 0 ['input_3[0][0]', 'reconstruction[0][0]'] ================================================================================================== Total params: 81365 (317.83 KB) Trainable params: 81365 (317.83 KB) Non-trainable params: 0 (0.00 Byte) __________________________________________________________________________________________________
最小复现代码
import os os.environ["TF_CPP_MIN_LOG_LEVEL"] = "3" import tensorflow as tf import numpy as np def generator(): for _ in range(1000): yield (np.random.random((240,120,1))*255).astype(np.uint16) def addLabels(x): return x,0 tfds = tf.data.Dataset.from_generator(generator, output_signature=tf.TensorSpec(shape=(240,120,1), dtype=tf.uint16)) tfds = tfds.map(addLabels) for x,y in tfds.take(1): print(x) print(y) from keras.optimizers import Adam from keras.layers import Input, MaxPooling2D, UpSampling2D, Layer from keras.models import Model class ScoreLayer(Layer): def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) def build(self, input_shape): return super().build(input_shape) def compute_output_shape(self, input_shape): return input_shape[0] @tf.function def call(self, x, y, *args, **kwargs): return tf.reduce_sum(tf.abs(x-y), name="diffScore")/1000 def build_autoencoder(): inputs = Input(shape=(240,120,1)) x = MaxPooling2D((4,4))(inputs) x = UpSampling2D((4,4))(x) x = ScoreLayer(name="score")(inputs, x) outputs = x autoencoder = Model(inputs, outputs) autoencoder.compile(optimizer=Adam(learning_rate=1e-2), loss='mae') return autoencoder model = build_autoencoder() model.summary() model.fit(tfds.batch(10).take(1))
问题原因
- ScoreLayer输出形状错误:
call方法中tf.reduce_sum默认对所有维度求和,导致输出是无维度的标量(形状()),而Keras的MAE损失期望模型输出带有批次维度(形状(batch_size,))。当批次数据输入时,模型输出没有维度,损失函数尝试访问y_pred[-1]时触发索引越界。 - compute_output_shape定义错误:返回的
input_shape[0]不符合实际输出形状,应该返回带批次维度的标量形状。 - 标签与输出形状不匹配:数据集的标签是标量0(批次后形状
(batch_size,)),但模型输出是无维度的标量,两者形状无法对齐。
解决方案
修正ScoreLayer的实现,确保输出保留批次维度,同时修正形状定义:
修正后的ScoreLayer
class ScoreLayer(Layer): def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) def build(self, input_shape): return super().build(input_shape) def compute_output_shape(self, input_shape): # 输入是两个(batch, 240, 120, 1)的张量,输出是(batch,) return (input_shape[0][0],) @tf.function def call(self, x, y, *args, **kwargs): # 仅对空间维度求和,保留批次维度 return tf.reduce_sum(tf.abs(x-y), axis=[1,2,3], name="diffScore")/1000
验证结果
修正后,模型摘要中ScoreLayer的输出形状会变为(None,),与标签形状(batch_size,)匹配,训练时不再触发索引越界错误,损失可正常计算。
内容的提问来源于stack exchange,提问作者TheClockTwister
相关产品推荐
相关产品推荐

