You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

堆叠LSTM模型推理阶段自定义预测方法实现求助

问题描述

用Keras构建了堆叠LSTM的多元时间序列一步预测模型,训练完成后尝试自定义前向传播替代内置predict方法,但无法得到正确结果。模型结构如下:

n_steps_in = 1
n_features = 3
model = Sequential()
model.add(LSTM(15, activation='tanh', return_sequences=True, input_shape=(n_steps_in,n_features)))
model.add(LSTM(20, activation='relu'))
model.add(Dense(n_features))
model.compile(optimizer='adam', loss='mse')
model.fit(X_train_lstm, y_train_lstm, epochs=100, validation_data = (X_test_lstm, y_test_lstm), 
            verbose=1)

自定义预测函数尝试实现单时间步预测(输入(1,1,3),输出(1,3)),但逻辑存在错误。


修正后的自定义LSTM预测方法

核心错误修正点

  • 层间输入传递错误:第二层LSTM的输入应为第一层LSTM的隐藏状态输出,而非原始输入
  • 维度处理错误:输入input_pred的时间步维度((1,1,3))需要压缩为(1,3),才能与权重矩阵正确运算
  • 细胞状态更新错误:门控与细胞状态的运算应为逐元素相乘,而非矩阵乘法
  • 变量名笔误:最终输出引用了未定义的layer3_output,应改为dense_output
  • 状态参数匹配:堆叠LSTM需要分别传入两层的初始细胞状态(c)和隐藏状态(h)

完整修正代码

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def model_predict_LSTM(model, input_pred, c0_tm1, h0_tm1, c1_tm1, h1_tm1):
    # 处理输入维度:(1,1,3) -> (1,3),去掉时间步维度
    x = input_pred.squeeze(axis=1)  # 形状变为(1, 3)
    
    # -------------------------- 第一层LSTM (units=15) --------------------------
    units0 = 15
    # 提取第一层权重
    W0, U0, b0 = model.layers[0].get_weights()
    
    # 拆分权重到四个门:输入门(i)、遗忘门(f)、候选细胞(c)、输出门(o)
    W0_i = W0[:, :units0]
    W0_f = W0[:, units0: units0*2]
    W0_c = W0[:, units0*2: units0*3]
    W0_o = W0[:, units0*3:]
    
    U0_i = U0[:, :units0]
    U0_f = U0[:, units0: units0*2]
    U0_c = U0[:, units0*2: units0*3]
    U0_o = U0[:, units0*3:]
    
    b0_i = b0[:units0]
    b0_f = b0[units0: units0*2]
    b0_c = b0[units0*2: units0*3]
    b0_o = b0[units0*3:]
    
    # 计算门控值
    i0_t = sigmoid(np.dot(x, W0_i) + np.dot(h0_tm1, U0_i) + b0_i)
    f0_t = sigmoid(np.dot(x, W0_f) + np.dot(h0_tm1, U0_f) + b0_f)
    o0_t = sigmoid(np.dot(x, W0_o) + np.dot(h0_tm1, U0_o) + b0_o)
    new_c0_t = np.tanh(np.dot(x, W0_c) + np.dot(h0_tm1, U0_c) + b0_c)
    
    # 更新细胞状态和隐藏状态(逐元素相乘)
    c0_t = f0_t * c0_tm1 + i0_t * new_c0_t
    h0_t = o0_t * np.tanh(c0_t)
    
    # -------------------------- 第二层LSTM (units=20) --------------------------
    units1 = 20
    # 提取第二层权重
    W1, U1, b1 = model.layers[1].get_weights()
    
    # 拆分权重到四个门
    W1_i = W1[:, :units1]
    W1_f = W1[:, units1: units1*2]
    W1_c = W1[:, units1*2: units1*3]
    W1_o = W1[:, units1*3:]
    
    U1_i = U1[:, :units1]
    U1_f = U1[:, units1: units1*2]
    U1_c = U1[:, units1*2: units1*3]
    U1_o = U1[:, units1*3:]
    
    b1_i = b1[:units1]
    b1_f = b1[units1: units1*2]
    b1_c = b1[units1*2: units1*3]
    b1_o = b1[units1*3:]
    
    # 第二层输入是第一层的隐藏状态h0_t(形状(1,15))
    i1_t = sigmoid(np.dot(h0_t, W1_i) + np.dot(h1_tm1, U1_i) + b1_i)
    f1_t = sigmoid(np.dot(h0_t, W1_f) + np.dot(h1_tm1, U1_f) + b1_f)
    o1_t = sigmoid(np.dot(h0_t, W1_o) + np.dot(h1_tm1, U1_o) + b1_o)
    new_c1_t = np.tanh(np.dot(h0_t, W1_c) + np.dot(h1_tm1, U1_c) + b1_c)
    
    # 更新细胞状态和隐藏状态(逐元素相乘)
    c1_t = f1_t * c1_tm1 + i1_t * new_c1_t
    h1_t = o1_t * np.tanh(c1_t)
    
    # -------------------------- Dense输出层 --------------------------
    dense_weights, dense_bias = model.layers[2].get_weights()
    dense_output = np.dot(h1_t, dense_weights) + dense_bias
    output_pred = dense_output.reshape(1, 3)
    
    # 返回预测结果和更新后的状态(用于多步预测)
    return output_pred, c0_t, h0_t, c1_t, h1_t

使用说明

  1. 初始状态初始化:首次预测时,初始状态可设为全零矩阵,维度匹配各层units:
    # 第一层初始状态:(1,15),第二层初始状态:(1,20)
    c0_init = np.zeros((1, 15))
    h0_init = np.zeros((1, 15))
    c1_init = np.zeros((1, 20))
    h1_init = np.zeros((1, 20))
    
  2. 单步预测调用:
    # 假设input_pred是形状为(1,1,3)的输入
    pred, c0_new, h0_new, c1_new, h1_new = model_predict_LSTM(model, input_pred, c0_init, h0_init, c1_init, h1_init)
    
  3. 多步预测:将上一步的c0_new, h0_new, c1_new, h1_new作为下一步的初始状态传入即可。

内容的提问来源于stack exchange,提问作者YIdirm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:55:38