You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LSTM训练时输入输出值ValueError(形状不兼容)问题排查

LSTM训练时出现形状不兼容报错

报错信息

训练基础LSTM网络时抛出如下错误:

Traceback (most recent call last):
  File "C:/Users/dell/Desktop/test run for LSTM thingy.py", line 39, in <module>
    history = model.fit(x_train, y_train, epochs=1, batch_size=16, verbose=1)
  File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "C:\Users\dell\AppData\Local\Temp\__autograph_generated_fileu1zdna1b.py", line 15, in tf__train_function
    retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope)
ValueError: in user code:

    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\training.py", line 1051, in train_function  *
        return step_function(self, iterator)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\training.py", line 1040, in step_function  **
        outputs = model.distribute_strategy.run(run_step, args=(data,))
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\training.py", line 1030, in run_step  **
        outputs = model.train_step(data)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\training.py", line 890, in train_step
        loss = self.compute_loss(x, y, y_pred, sample_weight)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\training.py", line 948, in compute_loss
        return self.compiled_loss(
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\compile_utils.py", line 201, in __call__
        loss_value = loss_obj(y_t, y_p, sample_weight=sw)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\losses.py", line 139, in __call__
        losses = call_fn(y_true, y_pred)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\losses.py", line 243, in call  **
        return ag_fn(y_true, y_pred, **self._fn_kwargs)
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\losses.py", line 1787, in categorical_crossentropy
        return backend.categorical_crossentropy(
    File "C:\Users\dell\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\backend.py", line 5119, in categorical_crossentropy
        target.shape.assert_is_compatible_with(output.shape)

    ValueError: Shapes (None, 133, 1320) and (None, 133, 5) are incompatible

待调试代码

import tensorflow as tf
x_train = tf.random.normal((28, 133, 1320))
y_train = tf.random.normal((28, 133, 1320))
model = tf.keras.Sequential()
model.add(tf.keras.layers.LSTM(5,activation='tanh',recurrent_activation='sigmoid', input_shape=(x_train.shape[1],x_train.shape[2]),return_sequences=True))
model.add(tf.keras.layers.Dense(5, activation= "softmax"))
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=0.0001), loss='categorical_crossentropy', metrics=['accuracy'])
model.summary()
history = model.fit(x_train, y_train, epochs=1, batch_size=16, verbose=1)

补充信息

  • 输入X形状:(28, 133, 1320)
  • 标签Y形状:(28, 133, 1320)
  • 任务目标:输出5个分类
  • 后续项目需使用X、Y形状结构相同的网络,初步判断问题与损失函数相关

问题根因

报错由三个不匹配问题共同导致:

  1. 输出维度和标签最后一维不匹配:模型开了return_sequences=True后每个时间步都会输出,接Dense(5)后最终输出形状是(批次大小, 133, 5),但传入的标签最后一维是1320,维度差直接触发形状兼容错误。
  2. 标签格式不符合分类任务要求:当前y_train是tf.random.normal生成的连续值,既不是分类任务要求的整数类别ID,也不是one-hot编码向量。
  3. 损失函数和标签格式不匹配:categorical_crossentropy要求输入标签为one-hot编码格式,连续值输入无法计算交叉熵损失。

修复方案

根据实际任务场景二选一调整即可:

场景1:逐时间步5分类(如序列标注、时序逐点分类任务)

保留输入输出时序长度一致的结构,修改标签和损失函数:

  • 标签调整为形状(28, 133),每个位置存储0-4的整数类别ID,对应每个时间步的分类结果;如果用one-hot格式则标签形状为(28, 133, 5)。
  • 整数标签搭配损失函数sparse_categorical_crossentropy,无需手动转one-hot;one-hot标签可保留原categorical_crossentropy。

可直接运行的修正代码:

import tensorflow as tf
x_train = tf.random.normal((28, 133, 1320))
# 生成合法的逐时间步分类标签
y_train = tf.random.uniform((28, 133), minval=0, maxval=5, dtype=tf.int32)

model = tf.keras.Sequential()
model.add(tf.keras.layers.LSTM(5,activation='tanh',recurrent_activation='sigmoid', 
                               input_shape=(x_train.shape[1],x_train.shape[2]),
                               return_sequences=True))
model.add(tf.keras.layers.Dense(5, activation= "softmax"))
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=0.0001), 
              loss='sparse_categorical_crossentropy', 
              metrics=['accuracy'])
model.summary()
history = model.fit(x_train, y_train, epochs=1, batch_size=16, verbose=1)

场景2:整序列单输出5分类(如序列分类任务)

调整模型输出结构,取消逐时间步输出:

  • 删除LSTM层的return_sequences=True参数(默认值为False),此时LSTM仅返回最后一个时间步的特征,模型最终输出形状为(批次大小, 5),符合单样本5分类的输出要求。
  • 标签调整为形状(28,)(存储0-4整数ID)或(28,5)(存储one-hot编码),对应搭配sparse_categorical_crossentropy或categorical_crossentropy损失即可。

内容的提问来源于stack exchange,提问作者chcheetah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 14:27:15