堆叠LSTM实现文本生成时遇activation_1维度不匹配错误求助
Hey Sam, let's break down what's causing that ValueError and fix it step by step.
What's Going Wrong?
Your error boils down to a mismatch between your model's output shape and your target data y's shape:
- Every LSTM layer in your code has
return_sequences=True. This means each layer outputs a full sequence of time steps, resulting in a 3D tensor shape:(number_of_samples, maxlen, number_of_features). - But your target data
yis 2D:(126984, 67), where each row represents a one-hot encoded prediction for the next single character in the sequence. - The final
Denselayer, when attached to a 3D input, will apply the dense transformation to every time step—keeping the output 3D. This clashes with your 2Dytensor, hence the error.
The Fix & Why It Works
The simple solution is to set return_sequences=False on your last LSTM layer:
return_sequences=Trueis only needed for stacking LSTM layers (so the next layer can process each time step of the previous layer's output).return_sequences=Falsetells the final LSTM layer to only output the last time step's features (a 2D tensor:(number_of_samples, 128)). This matches the input shape expectation of theDenselayer, which will then output a 2D tensor(number_of_samples, 67)—perfectly aligned with yourydata.
Also, let's fix those Keras version warnings: dropout_W is deprecated in Keras 2; use dropout instead.
Corrected Full Code
import numpy as np from keras.models import Sequential from keras.layers import LSTM, Dense, Activation import io path = 'lyrics.txt' with io.open(path, encoding='utf-8') as f: text = f.read().lower() print('corpus length:', len(text)) chars = sorted(list(set(text))) print('total chars:', len(chars)) char_indices = dict((c, i) for i, c in enumerate(chars)) indices_char = dict((i, c) for i, c in enumerate(chars)) # Cut text into semi-redundant sequences maxlen = 140 step = 3 sentences = [] next_chars = [] for i in range(0, len(text) - maxlen, step): sentences.append(text[i: i + maxlen]) next_chars.append(text[i + maxlen]) print('nb sequences:', len(sentences)) print('Vectorization...') x = np.zeros((len(sentences), maxlen, len(chars)), dtype=np.bool) y = np.zeros((len(sentences), len(chars)), dtype=np.bool) for i, sentence in enumerate(sentences): for t, char in enumerate(sentence): x[i, t, char_indices[char]] = 1 y[i, char_indices[next_chars[i]]] = 1 # Build stacked LSTM model print('Build model...') model = Sequential() # First two LSTMs return full sequences for stacking model.add(LSTM(128, dropout=0.5, return_sequences=True, input_shape=(maxlen, len(chars)))) model.add(LSTM(128, dropout=0.5, return_sequences=True)) # Final LSTM returns only the last time step's output model.add(LSTM(128, dropout=0.5, return_sequences=False)) model.add(Dense(len(chars))) model.add(Activation('softmax')) model.compile(loss='categorical_crossentropy', optimizer='adam')
Quick Extra Note
Since you're doing character-level next-character prediction, you only need the final LSTM's last time step output to predict the next character. If you were working on a sequence-to-sequence task (like machine translation), you'd want the final LSTM to return full sequences—but that's not the case here.
内容的提问来源于stack exchange,提问作者Sam B.

