You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras双输入字符串模型训练时出现形状不兼容错误求助

双输入Keras模型训练时形状不匹配问题排查与解决

问题场景

使用Keras构建双字符串输入的分类模型,模型预测功能正常,但执行训练时触发形状不兼容错误,错误提示为:

ValueError: Shapes (None, 1) and (None, 20) are incompatible

模型代码

vectorize_layer1 = TextVectorization(split="character", output_sequence_length=512,
                                     max_tokens=MAX_STRING_SIZE)
vectorize_layer1.adapt(list(vocab))

# 定义两组输入
inputA = Input(shape=(1,), dtype=tf.string)
inputB = Input(shape=(1,), dtype=tf.string)

# 第一个分支处理第一个输入
x = vectorize_layer1(inputA)
x = Embedding(len(vectorize_layer1.get_vocabulary()), MAX_STRING_SIZE)(x)
x = Bidirectional(LSTM(MAX_STRING_SIZE, return_sequences=True, dropout=.2))(x)
x = LSTM(MAX_STRING_SIZE, activation="tanh", return_sequences=False, dropout=.2)(x)
x = Model(inputs=inputA, outputs=x)

# 第二个分支处理第二个输入
y = vectorize_layer1(inputB)
y = Embedding(len(vectorize_layer1.get_vocabulary()), MAX_STRING_SIZE)(y)
y = Bidirectional(LSTM(MAX_STRING_SIZE, return_sequences=True, dropout=.2))(y)
y = LSTM(MAX_STRING_SIZE, activation="tanh", return_sequences=False, dropout=.2)(y)
y = Model(inputs=inputB, outputs=y)

# 合并两个分支的输出
combined = concatenate([x.output, y.output])
# 全连接层处理后输出分类结果
z = Dense(2, activation="relu")(combined)
z = Dense(len(LABELS), activation="softmax")(z)

# 定义完整模型
model = Model(inputs=[x.input, y.input], outputs=z)
print(model.predict((np.array(["i love python"]), np.array(["test"]))))
# 预测功能正常

model.summary()
plot_model(model, to_file="model.png", show_shapes=True, show_layer_names=True)
model.compile(optimizer='Adam',
              loss=CategoricalCrossentropy(from_logits=False),
              metrics=["categorical_accuracy"])

stopper = EarlyStopping(monitor='val_categorical_accuracy', patience=10)
checkpointer = ModelCheckpoint("model-best.tf", save_best_only=True)

model.fit(
    training,
    callbacks=[stopper, checkpointer],
    steps_per_epoch=2048,
    validation_data=validation,
    batch_size=8,
    epochs=epochs
)
model.save(output, save_format='tf')

完整错误栈

ValueError: in user code:

    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/training.py", line 1051, in train_function  *
        return step_function(self, iterator)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/training.py", line 1040, in step_function  **
        outputs = model.distribute_strategy.run(run_step, args=(data,))
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/training.py", line 1030, in run_step  **
        outputs = model.train_step(data)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/training.py", line 890, in train_step
        loss = self.compute_loss(x, y, y_pred, sample_weight)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/training.py", line 948, in compute_loss
        return self.compiled_loss(
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/engine/compile_utils.py", line 201, in __call__
        loss_value = loss_obj(y_t, y_p, sample_weight=sw)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/losses.py", line 139, in __call__
        losses = call_fn(y_true, y_pred)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/losses.py", line 243, in call  **
        return ag_fn(y_true, y_pred, **self._fn_kwargs)
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/losses.py", line 1787, in categorical_crossentropy
        return backend.categorical_crossentropy(
    File "/Users/**/code/venv/lib/python3.10/site-packages/keras/backend.py", line 5119, in categorical_crossentropy
        target.shape.assert_is_compatible_with(output.shape)

    ValueError: Shapes (None, 1) and (None, 20) are incompatible

训练数据示例

((array(['foo'], dtype='<U6'), array(['bar'], dtype='<U26')), array([0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]))

错误原因

  1. 标签维度不匹配:模型输出形状为(None, 20)(对应20类的one-hot预测结果),但训练时传入的标签被解析为(None,1),说明标签在数据加载过程中被错误降维。
  2. 训练参数冲突:model.fit中同时设置了steps_per_epoch和batch_size,如果training是tf.data.Dataset类型,这两个参数可能冲突,导致批次数据的解析异常,进而引发标签维度错误。

解决方法

1. 修正标签维度

确保训练/验证数据的标签是二维数组(形状为(样本数, 20)):

  • 如果是numpy数组形式,检查标签形状:
    print(training[1].shape)  # 预期输出 (N, 20),N为样本数量
    
    如果输出为(N,),说明标签被错误压缩,需重新整理为二维:
    # 假设原标签是由多个20维数组组成的列表/一维数组
    training_labels = np.array([label for label in training[1]])
    # 重新构建训练数据
    training = ((training[0][0], training[0][1]), training_labels)
    
  • 如果是tf.data.Dataset,确保在映射处理时不要使用squeeze()等压缩维度的操作,保留标签的二维结构。

2. 调整训练参数

如果training是已通过batch()方法设置批次大小的tf.data.Dataset,移除model.fit中的batch_size参数,避免参数冲突:

model.fit(
    training,
    callbacks=[stopper, checkpointer],
    # 若training是已分批的Dataset,移除batch_size
    # batch_size=8,
    validation_data=validation,
    epochs=epochs
)

同时,steps_per_epoch可设置为总样本数//批次大小,或直接省略,让Keras自动计算。

3. 验证数据结构

训练前手动检查一批数据的形状,确保匹配:

# 针对tf.data.Dataset
for inputs, labels in training.take(1):
    print("Input A形状:", inputs[0].shape)
    print("Input B形状:", inputs[1].shape)
    print("标签形状:", labels.shape)  # 必须为 (batch_size, 20)

内容的提问来源于stack exchange,提问作者reptilicus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 06:54:35