Keras中N单元LSTM与N个1单元并行LSTM是否等价及实现疑问
Great question—since you're just getting started with LSTMs for time series and digging into that convex basis LSTM paper, let's break down your questions clearly:
1. Is an LSTM layer with units=N equivalent to N separate LSTM layers each with units=1?
Short answer: No, they are not equivalent. Here's why:
- An
LSTM(units=N)layer contains N individual recurrent units, but all these units share a unified weight structure (the layer uses single sets of input-to-hidden and hidden-to-hidden weight matrices scaled to N units). - N separate
LSTM(units=1)layers each have their own independent weight matrices, biases, and internal cell states. This creates a drastically different parameter count:- For
LSTM(units=N), parameter count is4 * (input_dim * N + N² + N)(accounting for the 4 gates in every LSTM unit). - For N
LSTM(units=1)layers, parameter count isN * 4 * (input_dim * 1 + 1² + 1) = 4N*(input_dim + 2).
- For
- Additionally, the
LSTM(units=N)outputs a single tensor (shape(batch_size, timesteps, N)ifreturn_sequences=True, or(batch_size, N)if not), while N separate layers produce N distinct tensors.
2. Can an LSTM(units=N+1) replicate the concatenated output of LSTM(units=N) and LSTM(units=1)?
Theoretically, you might get a rough approximation with perfect weight tuning—but in practice, these are not equivalent. The LSTM(units=N+1) trains all N+1 units jointly, with their internal states and weights interacting as a single group. The two separate LSTMs (N and 1) train independently, creating a different parameter optimization landscape. You can't guarantee the combined layer will learn to split its output into the exact "y1 + y2" structure you're targeting.
3. How to implement parallel LSTMs in Keras?
You're spot-on—Sequential models can't handle parallel branches since they only support linear layer stacking. You'll need to use Keras' Functional API to build this architecture. Here's a practical example:
from tensorflow.keras.models import Model from tensorflow.keras.layers import Input, LSTM, Concatenate, Dense # Define input shape (adjust timesteps and input_dim to match your time series data) input_layer = Input(shape=(timesteps, input_dim)) # Create parallel LSTM branches lstm_branch_n = LSTM(units=N, return_sequences=False)(input_layer) lstm_branch_1 = LSTM(units=1, return_sequences=False)(input_layer) # Merge outputs of the two branches merged_output = Concatenate()([lstm_branch_n, lstm_branch_1]) # Add downstream layers (e.g., a dense layer for final prediction) output_layer = Dense(units=1)(merged_output) # Build and compile the model model = Model(inputs=input_layer, outputs=output_layer) model.compile(optimizer='adam', loss='mse')
If you need to preserve the sequence dimension for further processing (like another LSTM layer), set return_sequences=True in both parallel LSTM layers. The Concatenate layer will stack outputs along the last dimension, resulting in a tensor of shape (batch_size, timesteps, N+1).
内容的提问来源于stack exchange,提问作者Leo

