You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Keras中LSTM工作机制的技术问询及输出形状验证

Understanding LSTM Output Shape and Logic in Keras (With Translation Task Example)

Let's break down exactly why your model outputs a shape of (3,1) and walk through how the LSTM processes your input step-by-step.

First, Clarify the Key Parameters

Your input is a sequence of 3 words, each represented as a 4-dimensional vector — so the input shape (ignoring batch size) is (3,4). Here's what each part of your LSTM layer does:

  • LSTM(1): The 1 defines the dimension of the LSTM's hidden state output. Every time the LSTM processes a time step, it outputs a 1-dimensional vector representing its current state.
  • return_sequences=True: This is the critical setting here. When enabled, the LSTM doesn't just return the hidden state from the final time step — it returns the hidden state generated at every single time step in the input sequence.

Why the Output Shape is (3,1)

Let's map this directly to your example sequence (how are you, 3 time steps):

  1. Time Step 1: The LSTM takes the first word's 4D vector, processes it, and outputs a 1D hidden state (h1).
  2. Time Step 2: The LSTM takes the second word's 4D vector plus the hidden state h1 from the previous step, processes them together, and outputs a new 1D hidden state (h2).
  3. Time Step 3: The LSTM takes the third word's 4D vector plus h2, processes them, and outputs h3.

Since return_sequences=True, the model stacks these three 1D states into a sequence: [h1, h2, h3], which gives you a shape of (3,1) (note: Keras actually shows (None, 3,1) in model.summary() — the None represents a variable batch size, but when feeding a single sample, it simplifies to (3,1)).

How This Fits Into a Translation Task

In translation scenarios, this setup makes sense if you're doing word-level sequence-to-sequence mapping (e.g., generating a translated word for every input word). If you instead wanted a single "summary" vector of the entire input sequence to feed into a decoder LSTM (the standard encoder-decoder translation setup), you'd set return_sequences=False (the default). That would give you only the final hidden state h3, with a shape of (1) (or (batch_size,1) for batches).

Quick Test: Adjusting return_sequences

If you modify your code to:

model.add(LSTM(1, input_shape=(3,4)))  # return_sequences=False by default

Your output shape would change to (1) (or (batch_size,1)), since only the final time step's hidden state is returned.


内容的提问来源于stack exchange,提问作者Pusheen_the_dev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:23:58