LSTM模型训练报错求助:预期3维目标却得到形状(1,184800)数组
Let's walk through what's causing this error and fix your LSTM training code step by step—this is a common gotcha with sequence models in Keras.
What's Going Wrong?
Target Shape Mismatch
Your final LSTM layer usesreturn_sequences=True, which means it outputs a 3D tensor (batch_size, sequence_length, num_units). But yourY_trainis reshaped to a 2D array(1, 184800)—Keras can't match these dimensions, hence the error.Incorrect Data Reshaping
You're treating your entire training dataset as one giant sequence (X_train.reshape(1, n_steps, n_features)). That's almost never the right approach unless your data is a single continuous time series (and even then, you'd split it into smaller chunks). LSTMs expect inputs in the shape(number_of_samples, timesteps_per_sample, number_of_features)—each row in your CSV should be a sample, not a timestep in one big sequence.Loss Function Misalignment
If you're predicting a continuous value (yourYlooks like it's from a CSV column of floats), usingsparse_categorical_crossentropyis wrong. That loss is for multi-class classification tasks, not regression.
Fixed Code Examples
Let's cover two common scenarios based on what you're trying to do:
Scenario 1: Single-Step Prediction (Most Likely Your Use Case)
If each row in your dataset is a sample with 25 features, and you want to predict one corresponding target value:
import pandas as pd import numpy as np from sklearn.model_selection import train_test_split from keras.models import Sequential from keras.layers import Dense from keras.layers import LSTM # Load and prepare data dataframe = pd.read_csv("AHPS_26C_Norm.csv", header=None) dataset = dataframe.values X = dataset[1:, 1:26].astype(float) Y = dataset[1:, 0].astype(float) # Ensure Y is float for regression X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.20, random_state=42) # Reshape inputs for LSTM: (samples, timesteps, features) # Here, each sample is a 1-timestep sequence (since each row is independent) n_features = 25 X_train = X_train.reshape(X_train.shape[0], 1, n_features) X_test = X_test.reshape(X_test.shape[0], 1, n_features) # Reshape targets to 2D (samples, 1) for single output Y_train = Y_train.reshape(Y_train.shape[0], 1) Y_test = Y_test.reshape(Y_test.shape[0], 1) # Build the model model = Sequential() model.add(LSTM(100, activation='relu', input_shape=(1, n_features))) model.add(Dense(1)) # Single output value per sample model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae']) # Train the model model.fit(X_train, Y_train, epochs=50, validation_data=(X_test, Y_test), verbose=1) # Evaluate scores = model.evaluate(X_train, Y_train, verbose=0) print(f"Loss: {scores[0]:.4f}, MAE: {scores[1]:.4f}")
Scenario 2: Sequence-to-Sequence Prediction
If you actually need to predict a sequence of outputs matching the length of your input sequence (e.g., time series forecasting where you predict multiple future steps):
import pandas as pd import numpy as np from sklearn.model_selection import train_test_split from keras.models import Sequential from keras.layers import Dense from keras.layers import LSTM from keras.layers import RepeatVector from keras.layers import TimeDistributed # Load data (same as before) dataframe = pd.read_csv("AHPS_26C_Norm.csv", header=None) dataset = dataframe.values X = dataset[1:, 1:26].astype(float) Y = dataset[1:, 0].astype(float) # For sequence-to-sequence, we'll treat the entire dataset as one sequence (adjust if needed) X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.20, random_state=42) n_steps = X_train.shape[0] n_features = 25 # Reshape input to (1, n_steps, n_features) (one sample, full sequence) X_train = X_train.reshape(1, n_steps, n_features) X_test = X_test.reshape(1, X_test.shape[0], n_features) # Reshape target to 3D: (1, n_steps, 1) to match model's 3D output Y_train = Y_train.reshape(1, n_steps, 1) Y_test = Y_test.reshape(1, X_test.shape[1], 1) # Build sequence-to-sequence model model = Sequential() model.add(LSTM(100, activation='relu', batch_input_shape=(1, n_steps, n_features))) model.add(RepeatVector(n_steps)) model.add(LSTM(100, activation='relu', return_sequences=True)) model.add(TimeDistributed(Dense(1))) # Output one value per timestep model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae']) # Train model.fit(X_train, Y_train, epochs=50, validation_data=(X_test, Y_test), batch_size=1, verbose=1) # Evaluate scores = model.evaluate(X_train, Y_train, verbose=0) print(f"Loss: {scores[0]:.4f}, MAE: {scores[1]:.4f}")
Key Notes
- I reduced the epochs from 300 to 50 to avoid overfitting during testing—you can adjust this back once the code runs.
- For classification tasks (if
Yis integer labels), swap the loss function tosparse_categorical_crossentropyand change the final Dense layer toDense(num_classes, activation='softmax'). - Avoid hardcoding
batch_size=1unless you have a specific reason (like memory constraints)—Keras handles batch sizing automatically by default.
内容的提问来源于stack exchange,提问作者Iliya Juhari

