LSTM输入维度错误求助:销售预测模型报(0,1)形状数组异常
Hey Dan, let's break down why you're hitting this error and how to fix it. The core issue is your model is receiving an empty or incorrectly shaped input array, which triggers the dimension mismatch. Here's what's going wrong and how to resolve each problem:
1. Your Dataset is Too Small for n_prev=100
The biggest culprit is likely the n_prev=100 parameter in your _load_data function. This tells the code to use 100 past time steps to predict the next one. If your CSV has fewer than 100 rows (or even fewer once split into train/test), the loop in _load_data won’t run at all—resulting in empty X_train and X_test arrays (shape (0, ...)).
Fix:
- First, check how much data you actually have by adding a print statement after loading the CSV:
data = read_csv(path, header=0, index_col=0, usecols=COLS) print(f"Total rows in dataset: {len(data)}") # Confirm this is >= n_prev - If your dataset is smaller than 100 rows, reduce
n_prevto a realistic value. For example, if you have 50 rows, tryn_prev=10:def _load_data(df, n_prev=10): # Lower this to match your dataset size docX, docY = [], [] # Guard against invalid dataset size if len(df) <= n_prev: raise ValueError(f"Dataset length ({len(df)}) must be larger than n_prev ({n_prev})") for i in range(len(df) - n_prev): docX.append(df.iloc[i:i + n_prev].values) docY.append(df.iloc[i + n_prev].values) alsX = np.array(docX) alsY = np.array(docY) return alsX, alsY
2. Validate Input Shape for LSTM
LSTMs require 3-dimensional input: (number_of_samples, timesteps, number_of_features). Your code should produce this if n_prev is set correctly, but let’s confirm and adjust the model to be explicit:
Fix:
- After splitting your data, print the shapes to verify:
X_train, y_train, X_test, y_test = train_test_split(data) print(f"X_train shape: {X_train.shape}") # Should look like (samples, n_prev, 6) print(f"y_train shape: {y_train.shape}") # Should look like (samples, 6) - Update your LSTM layer to use the exact timesteps from
n_prev(explicit is better than implicit here):model.add(LSTM(49, input_shape=(n_prev, data.shape[1]))) # Use dynamic feature count - Remove the
input_dim=49from your Dense layer—it’s redundant (the layer automatically uses the output size of the LSTM):model.add(Dense(data.shape[1])) # No need for input_dim here
3. Fix Data Loading Issues
Your usecols=[0,1,2,3,4,5,6] skips the 7th column (Sunday, since you have 8 total columns). If you meant to include all days, adjust COLS to:
COLS = [0,1,2,3,4,5,6,7] # Includes weekend date + all 7 days
Also, double-check your file path is correct and that the CSV has a valid header row (since you set header=0).
Corrected Full Code Example
Here’s the adjusted code incorporating all fixes (assuming n_prev=10 based on a smaller dataset):
import numpy as np from pandas import read_csv from keras.models import Sequential from keras.layers.core import Dense from keras.layers.recurrent import LSTM path = '/Users/dan/PycharmProjects/test1/sales_EXP.csv' COLS = [0,1,2,3,4,5,6,7] # Include all 8 columns n_prev = 10 # Adjust based on your dataset size # Load and verify dataset data = read_csv(path, header=0, index_col=0, usecols=COLS) print(f"Loaded dataset shape: {data.shape}") def _load_data(df, n_prev): docX, docY = [], [] if len(df) <= n_prev: raise ValueError(f"Dataset length ({len(df)}) must be greater than n_prev ({n_prev})") for i in range(len(df) - n_prev): docX.append(df.iloc[i:i + n_prev].values) docY.append(df.iloc[i + n_prev].values) return np.array(docX), np.array(docY) def train_test_split(df, test_size=0.1): num_train = round(len(df) * (1 - test_size)) if num_train <= n_prev: raise ValueError(f"Training subset size ({num_train}) must be greater than n_prev ({n_prev})") X_train, y_train = _load_data(df.iloc[0:num_train], n_prev) X_test, y_test = _load_data(df.iloc[num_train:], n_prev) return X_train, y_train, X_test, y_test # Build model model = Sequential() model.add(LSTM(49, input_shape=(n_prev, data.shape[1]))) model.add(Dense(data.shape[1])) model.compile(loss="mse", optimizer="adam") # Split data and train X_train, y_train, X_test, y_test = train_test_split(data) print(f"X_train shape: {X_train.shape}, y_train shape: {y_train.shape}") # Adjust batch size to match your dataset (32 is a safe starting point) model.fit(X_train, y_train, batch_size=32, epochs=10, validation_split=0.05) # Evaluate predictions predicted = model.predict(X_test) rmse = np.sqrt(((predicted - y_test) ** 2).mean(axis=0)) print(f"\nRMSE per day: {np.around(rmse, 2)}")
Final Tips
- Start with fewer epochs (like 10) and adjust batch size to fit your dataset size—batch size 450 is way too large for small datasets.
- If you still get errors, double-check that
X_trainisn’t empty—this is almost always the root cause of the(0,1)shape error.
内容的提问来源于stack exchange,提问作者Dan

