You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

堆叠LSTM实现文本生成时遇activation_1维度不匹配错误求助

Fixing the Dimension Mismatch Error in Your LSTM Text Generation Model

Hey Sam, let's break down what's causing that ValueError and fix it step by step.

What's Going Wrong?

Your error boils down to a mismatch between your model's output shape and your target data y's shape:

  • Every LSTM layer in your code has return_sequences=True. This means each layer outputs a full sequence of time steps, resulting in a 3D tensor shape: (number_of_samples, maxlen, number_of_features).
  • But your target data y is 2D: (126984, 67), where each row represents a one-hot encoded prediction for the next single character in the sequence.
  • The final Dense layer, when attached to a 3D input, will apply the dense transformation to every time step—keeping the output 3D. This clashes with your 2D y tensor, hence the error.

The Fix & Why It Works

The simple solution is to set return_sequences=False on your last LSTM layer:

  • return_sequences=True is only needed for stacking LSTM layers (so the next layer can process each time step of the previous layer's output).
  • return_sequences=False tells the final LSTM layer to only output the last time step's features (a 2D tensor: (number_of_samples, 128)). This matches the input shape expectation of the Dense layer, which will then output a 2D tensor (number_of_samples, 67)—perfectly aligned with your y data.

Also, let's fix those Keras version warnings: dropout_W is deprecated in Keras 2; use dropout instead.

Corrected Full Code

import numpy as np
from keras.models import Sequential
from keras.layers import LSTM, Dense, Activation
import io

path = 'lyrics.txt'
with io.open(path, encoding='utf-8') as f:
    text = f.read().lower()
print('corpus length:', len(text))

chars = sorted(list(set(text)))
print('total chars:', len(chars))
char_indices = dict((c, i) for i, c in enumerate(chars))
indices_char = dict((i, c) for i, c in enumerate(chars))

# Cut text into semi-redundant sequences
maxlen = 140
step = 3
sentences = []
next_chars = []
for i in range(0, len(text) - maxlen, step):
    sentences.append(text[i: i + maxlen])
    next_chars.append(text[i + maxlen])
print('nb sequences:', len(sentences))

print('Vectorization...')
x = np.zeros((len(sentences), maxlen, len(chars)), dtype=np.bool)
y = np.zeros((len(sentences), len(chars)), dtype=np.bool)
for i, sentence in enumerate(sentences):
    for t, char in enumerate(sentence):
        x[i, t, char_indices[char]] = 1
    y[i, char_indices[next_chars[i]]] = 1

# Build stacked LSTM model
print('Build model...')
model = Sequential()
# First two LSTMs return full sequences for stacking
model.add(LSTM(128, dropout=0.5, return_sequences=True, input_shape=(maxlen, len(chars))))
model.add(LSTM(128, dropout=0.5, return_sequences=True))
# Final LSTM returns only the last time step's output
model.add(LSTM(128, dropout=0.5, return_sequences=False))
model.add(Dense(len(chars)))
model.add(Activation('softmax'))

model.compile(loss='categorical_crossentropy', optimizer='adam')

Quick Extra Note

Since you're doing character-level next-character prediction, you only need the final LSTM's last time step output to predict the next character. If you were working on a sequence-to-sequence task (like machine translation), you'd want the final LSTM to return full sequences—but that's not the case here.

内容的提问来源于stack exchange,提问作者Sam B.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:04:20