You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LSTM模型训练报错求助:预期3维目标却得到形状(1,184800)数组

Fixing ValueError: Expected 3D array for LSTM target but got 2D shape (1, 184800)

Let's walk through what's causing this error and fix your LSTM training code step by step—this is a common gotcha with sequence models in Keras.

What's Going Wrong?

  1. Target Shape Mismatch
    Your final LSTM layer uses return_sequences=True, which means it outputs a 3D tensor (batch_size, sequence_length, num_units). But your Y_train is reshaped to a 2D array (1, 184800)—Keras can't match these dimensions, hence the error.

  2. Incorrect Data Reshaping
    You're treating your entire training dataset as one giant sequence (X_train.reshape(1, n_steps, n_features)). That's almost never the right approach unless your data is a single continuous time series (and even then, you'd split it into smaller chunks). LSTMs expect inputs in the shape (number_of_samples, timesteps_per_sample, number_of_features)—each row in your CSV should be a sample, not a timestep in one big sequence.

  3. Loss Function Misalignment
    If you're predicting a continuous value (your Y looks like it's from a CSV column of floats), using sparse_categorical_crossentropy is wrong. That loss is for multi-class classification tasks, not regression.

Fixed Code Examples

Let's cover two common scenarios based on what you're trying to do:

Scenario 1: Single-Step Prediction (Most Likely Your Use Case)

If each row in your dataset is a sample with 25 features, and you want to predict one corresponding target value:

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from keras.models import Sequential
from keras.layers import Dense
from keras.layers import LSTM

# Load and prepare data
dataframe = pd.read_csv("AHPS_26C_Norm.csv", header=None)
dataset = dataframe.values
X = dataset[1:, 1:26].astype(float)
Y = dataset[1:, 0].astype(float)  # Ensure Y is float for regression

X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.20, random_state=42)

# Reshape inputs for LSTM: (samples, timesteps, features)
# Here, each sample is a 1-timestep sequence (since each row is independent)
n_features = 25
X_train = X_train.reshape(X_train.shape[0], 1, n_features)
X_test = X_test.reshape(X_test.shape[0], 1, n_features)

# Reshape targets to 2D (samples, 1) for single output
Y_train = Y_train.reshape(Y_train.shape[0], 1)
Y_test = Y_test.reshape(Y_test.shape[0], 1)

# Build the model
model = Sequential()
model.add(LSTM(100, activation='relu', input_shape=(1, n_features)))
model.add(Dense(1))  # Single output value per sample
model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae'])

# Train the model
model.fit(X_train, Y_train, epochs=50, validation_data=(X_test, Y_test), verbose=1)

# Evaluate
scores = model.evaluate(X_train, Y_train, verbose=0)
print(f"Loss: {scores[0]:.4f}, MAE: {scores[1]:.4f}")

Scenario 2: Sequence-to-Sequence Prediction

If you actually need to predict a sequence of outputs matching the length of your input sequence (e.g., time series forecasting where you predict multiple future steps):

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from keras.models import Sequential
from keras.layers import Dense
from keras.layers import LSTM
from keras.layers import RepeatVector
from keras.layers import TimeDistributed

# Load data (same as before)
dataframe = pd.read_csv("AHPS_26C_Norm.csv", header=None)
dataset = dataframe.values
X = dataset[1:, 1:26].astype(float)
Y = dataset[1:, 0].astype(float)

# For sequence-to-sequence, we'll treat the entire dataset as one sequence (adjust if needed)
X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.20, random_state=42)
n_steps = X_train.shape[0]
n_features = 25

# Reshape input to (1, n_steps, n_features) (one sample, full sequence)
X_train = X_train.reshape(1, n_steps, n_features)
X_test = X_test.reshape(1, X_test.shape[0], n_features)

# Reshape target to 3D: (1, n_steps, 1) to match model's 3D output
Y_train = Y_train.reshape(1, n_steps, 1)
Y_test = Y_test.reshape(1, X_test.shape[1], 1)

# Build sequence-to-sequence model
model = Sequential()
model.add(LSTM(100, activation='relu', batch_input_shape=(1, n_steps, n_features)))
model.add(RepeatVector(n_steps))
model.add(LSTM(100, activation='relu', return_sequences=True))
model.add(TimeDistributed(Dense(1)))  # Output one value per timestep
model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae'])

# Train
model.fit(X_train, Y_train, epochs=50, validation_data=(X_test, Y_test), batch_size=1, verbose=1)

# Evaluate
scores = model.evaluate(X_train, Y_train, verbose=0)
print(f"Loss: {scores[0]:.4f}, MAE: {scores[1]:.4f}")

Key Notes

  • I reduced the epochs from 300 to 50 to avoid overfitting during testing—you can adjust this back once the code runs.
  • For classification tasks (if Y is integer labels), swap the loss function to sparse_categorical_crossentropy and change the final Dense layer to Dense(num_classes, activation='softmax').
  • Avoid hardcoding batch_size=1 unless you have a specific reason (like memory constraints)—Keras handles batch sizing automatically by default.

内容的提问来源于stack exchange,提问作者Iliya Juhari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:15:05