You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras多分类输出模型的y_train数据格式优化求助

多输出Keras模型训练数据构建方案

你的模型属于多任务输出架构:一个分支做意图分类(4个类别),另一个分支做序列标注(每个单词对应10个标签),当前的y_train格式完全不符合模型需求,正确的构建方式需要拆分两个任务的标签数据,具体步骤如下:

1. 拆分两类标签数据

意图分类标签处理

假设你的数据中意图标签列名为intent,需要将类别转为模型可识别的格式:

from tensorflow.keras.utils import to_categorical

# 先把文本意图转为整数编码
intent_codes = df_ic['intent'].astype('category').cat.codes
# 转为one-hot编码,最终形状为 (样本数量, 4)
y_intent = to_categorical(intent_codes, num_classes=4)

序列标注标签处理

假设每个句子的单词标签存在token_labels列(每个元素是对应句子单词的标签列表,比如[2,5,0,...]),需要先对齐序列长度(和你输入文本的序列长度一致,比如你代码里的85),再转格式:

from tensorflow.keras.preprocessing.sequence import pad_sequences

# 提取所有句子的标签序列
token_label_list = df_ic['token_labels'].tolist()
# 对齐序列长度到85,短序列补0,长序列截断
padded_labels = pad_sequences(token_label_list, maxlen=85, padding='post', truncating='post')
# 转为one-hot编码,最终形状为 (样本数量, 85, 10)
y_token = to_categorical(padded_labels, num_classes=10)

2. 传入model.fit的正确格式

如果你的模型输出层有命名(比如分别叫intent_output和token_output),用字典传入更清晰:

model.fit(
    x_train,  # 预处理好的输入文本序列,形状为(样本数量, 85)
    {'intent_output': y_intent, 'token_output': y_token},
    epochs=10,
    batch_size=32
)

如果输出层没有命名,直接按模型输出顺序用列表传入:

model.fit(
    x_train,
    [y_intent, y_token],
    epochs=10,
    batch_size=32
)

关键注意事项

  • 输入文本x必须做对应预处理:分词、转整数序列、padding到85长度,确保和序列标注的输入维度匹配
  • 损失函数要对应两个输出分别设置:意图分类用CategoricalCrossentropy,序列标注也用CategoricalCrossentropy;如果不想做one-hot编码,也可以用整数标签配合SparseCategoricalCrossentropy,能节省内存

内容的提问来源于stack exchange,提问作者Martin Horvath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 02:45:10