You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras中N单元LSTM与N个1单元并行LSTM是否等价及实现疑问

Understanding Parallel LSTMs in Keras

Great question—since you're just getting started with LSTMs for time series and digging into that convex basis LSTM paper, let's break down your questions clearly:

1. Is an LSTM layer with units=N equivalent to N separate LSTM layers each with units=1?

Short answer: No, they are not equivalent. Here's why:

  • An LSTM(units=N) layer contains N individual recurrent units, but all these units share a unified weight structure (the layer uses single sets of input-to-hidden and hidden-to-hidden weight matrices scaled to N units).
  • N separate LSTM(units=1) layers each have their own independent weight matrices, biases, and internal cell states. This creates a drastically different parameter count:
    • For LSTM(units=N), parameter count is 4 * (input_dim * N + N² + N) (accounting for the 4 gates in every LSTM unit).
    • For N LSTM(units=1) layers, parameter count is N * 4 * (input_dim * 1 + 1² + 1) = 4N*(input_dim + 2).
  • Additionally, the LSTM(units=N) outputs a single tensor (shape (batch_size, timesteps, N) if return_sequences=True, or (batch_size, N) if not), while N separate layers produce N distinct tensors.

2. Can an LSTM(units=N+1) replicate the concatenated output of LSTM(units=N) and LSTM(units=1)?

Theoretically, you might get a rough approximation with perfect weight tuning—but in practice, these are not equivalent. The LSTM(units=N+1) trains all N+1 units jointly, with their internal states and weights interacting as a single group. The two separate LSTMs (N and 1) train independently, creating a different parameter optimization landscape. You can't guarantee the combined layer will learn to split its output into the exact "y1 + y2" structure you're targeting.

3. How to implement parallel LSTMs in Keras?

You're spot-on—Sequential models can't handle parallel branches since they only support linear layer stacking. You'll need to use Keras' Functional API to build this architecture. Here's a practical example:

from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, LSTM, Concatenate, Dense

# Define input shape (adjust timesteps and input_dim to match your time series data)
input_layer = Input(shape=(timesteps, input_dim))

# Create parallel LSTM branches
lstm_branch_n = LSTM(units=N, return_sequences=False)(input_layer)
lstm_branch_1 = LSTM(units=1, return_sequences=False)(input_layer)

# Merge outputs of the two branches
merged_output = Concatenate()([lstm_branch_n, lstm_branch_1])

# Add downstream layers (e.g., a dense layer for final prediction)
output_layer = Dense(units=1)(merged_output)

# Build and compile the model
model = Model(inputs=input_layer, outputs=output_layer)
model.compile(optimizer='adam', loss='mse')

If you need to preserve the sequence dimension for further processing (like another LSTM layer), set return_sequences=True in both parallel LSTM layers. The Concatenate layer will stack outputs along the last dimension, resulting in a tensor of shape (batch_size, timesteps, N+1).

内容的提问来源于stack exchange,提问作者Leo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:51:06