You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras多对一模型如何输出序列每个事件的预测得分?

问题描述

我有一个基于Keras训练的序列数据模型,每个序列对应单个标签。输入为经过类别编码的特征,先通过Embedding层再传入GRU层,模型代码如下:

samples, timesteps, features = 2000, 10, 1
inputs_1 = np.random.randint(1, 50, [samples, timesteps, features]).astype(np.float32)
labels = np.random.randint(0, 2, [samples, 1])

# Input
input_ = Input(shape=(None,))

# Embeddings
emb = Embedding(input_dim=int(50),
                            output_dim=20,
                            input_length=(None,),
                            mask_zero=False,
                            name="cat_feat_0" + "_emb")(input_)

gru = GRU(32,
          activation="tanh",
          dropout=0,
          recurrent_dropout=0,
          go_backwards=False,
          return_sequences=False,
          name="gru_cat")(emb)

y = Dense(10, activation = "tanh")(gru)
y = Dropout(0.4)(y)
y = Dense(1, activation = "sigmoid")(y)

model = Model(inputs=input_, outputs=y)
model.compile(loss=BCE_Last_Event,
                          optimizer=Adam(beta_1=0.9, beta_2=0.999),
                          metrics=["accuracy"])

model.predict(inputs_1).shape

当前预测输出形状为(2000,1),即每个序列对应一个标签。我希望模型能输出序列中每个事件的得分,使预测形状变为(2000,10,1)。我知道可以将GRU层设置为return_sequences=True来返回序列,但由于只有单个标签,损失函数会出现问题。我目前考虑两种方案:

  • 创建一个使用相同训练权重的新模型,返回序列预测结果;
  • 用TimeDistributed层包裹模型,实现每个事件的预测。

但我担心第二种方案仅以单个事件作为输入,而非整个序列,这种顾虑是否正确?最优解决方案是什么?


解答

你的顾虑完全正确:TimeDistributed如果直接包裹原模型,会把每个时间步的特征单独传入模型,完全丢失序列的时序依赖关系——原模型的GRU是基于整个序列训练的,拆成单个时间步输入后,根本无法利用上下文信息,所以这个方案不可行。

最优方案是你提到的第一种:基于训练好的原模型权重,构建一个专用的推理模型,让GRU返回全序列的隐藏状态,再通过TimeDistributed包装的全连接层生成每个时间步的得分。具体步骤如下:

1. 核心思路

原模型的GRU已经学习到了序列的时序特征,我们只需要修改GRU的return_sequences参数为True,让它输出每个时间步的隐藏状态,再用TimeDistributed把原模型的全连接层应用到每个时间步的隐藏状态上,就能得到每个事件的得分,同时完全复用训练好的权重。

2. 实现代码

# 假设原模型已完成训练,直接复用已有层和权重
# 1. 复用输入层和Embedding层
input_infer = input_  # 原模型的输入层
emb_infer = model.get_layer("cat_feat_0_emb")(input_infer)

# 2. 创建return_sequences=True的GRU层,并加载原GRU的权重
gru_infer = GRU(32,
                activation="tanh",
                dropout=0,
                recurrent_dropout=0,
                go_backwards=False,
                return_sequences=True,  # 修改为返回全序列
                name="gru_cat_infer")(emb_infer)
# 加载原模型GRU层的权重
gru_infer.set_weights(model.get_layer("gru_cat").get_weights())

# 3. 用TimeDistributed包装全连接层,复用原模型的Dense权重
# 获取原模型中全连接层的权重
dense1_weights = model.layers[3].get_weights()  # 对应Dense(10, tanh)
dense2_weights = model.layers[5].get_weights()  # 对应Dense(1, sigmoid)

# 构建时序全连接分支
y_infer = TimeDistributed(Dense(10, activation="tanh"))(gru_infer)
y_infer = TimeDistributed(Dropout(0.4))(y_infer)
y_infer = TimeDistributed(Dense(1, activation="sigmoid"))(y_infer)

# 给新的全连接层加载原权重
y_infer.layers[0].set_weights(dense1_weights)
y_infer.layers[2].set_weights(dense2_weights)

# 构建最终的推理模型
infer_model = Model(inputs=input_infer, outputs=y_infer)

# 验证输出形状
print(infer_model.predict(inputs_1).shape)  # 输出 (2000, 10, 1)

3. 关键说明

  • 无需重新训练:我们只是复用原模型训练好的特征提取能力,生成每个时间步的得分,完全不需要调整损失函数或重新训练;
  • 保留时序依赖:修改后的GRU返回每个时间步的隐藏状态,这些状态包含了从序列开头到当前步的上下文信息,确保每个事件的得分是基于整个序列的;
  • 权重复用:所有层的参数都直接从原模型加载,保证推理结果和原模型的全局预测逻辑一致。

内容的提问来源于stack exchange,提问作者Olivier

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 19:25:19