You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras构建有状态LSTM时batch_input_shape参数报错的解决咨询

问题:Keras有状态LSTM指定batch_input_shape时触发ValueError

错误信息

ValueError: Unrecognized keyword arguments passed to LSTM: {'batch_input_shape': (1, 1, 14)}

复现代码

import pandas as pd
from sklearn.preprocessing import MinMaxScaler
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, LSTM

# Load your data
file_path = 'path_to_your_file.csv'
data = pd.read_csv(file_path)

# Create 'date' column with the first day of each month
data['date'] = pd.to_datetime(data['tahun'].astype(str) + '-' + data['bulan'].astype(str) + '-01')
data['date'] = data['date'] + pd.offsets.MonthEnd(0)
data.set_index('date', inplace=True)

# Group by 'date' and sum the 'amaun_rm' column
df_sales = data.groupby('date')['amaun_rm'].sum().reset_index()

# Create a new dataframe to model the difference
df_diff = df_sales.copy()
df_diff['prev_amaun_rm'] = df_diff['amaun_rm'].shift(1)
df_diff = df_diff.dropna()
df_diff['diff'] = df_diff['amaun_rm'] - df_diff['prev_amaun_rm']

# Create new dataframe from transformation from time series to supervised
df_supervised = df_diff.drop(['prev_amaun_rm'], axis=1)
for inc in range(1, 13):
    field_name = 'lag_' + str(inc)
    df_supervised[field_name] = df_supervised['diff'].shift(inc)

# Adding moving averages
df_supervised['ma_3'] = df_supervised['amaun_rm'].rolling(window=3).mean().shift(1)
df_supervised['ma_6'] = df_supervised['amaun_rm'].rolling(window=6).mean().shift(1)
df_supervised['ma_12'] = df_supervised['amaun_rm'].rolling(window=12).mean().shift(1)
df_supervised = df_supervised.dropna().reset_index(drop=True)
df_supervised = df_supervised.fillna(df_supervised.mean())

# Split the data into train and test sets
train_set, test_set = df_supervised[0:-6].values, df_supervised[-6:].values
scaler = MinMaxScaler(feature_range=(-1, 1))
scaler = scaler.fit(train_set)
train_set_scaled = scaler.transform(train_set)
test_set_scaled = scaler.transform(test_set)

# Split into input and output
X_train, y_train = train_set_scaled[:, 1:], train_set_scaled[:, 0]
X_test, y_test = test_set_scaled[:, 1:], test_set_scaled[:, 0]
X_train = X_train.reshape((X_train.shape[0], 1, X_train.shape[1]))
X_test = X_test.reshape((X_test.shape[0], 1, X_test.shape[1]))

# Check the shape of X_train
print("X_train shape:", X_train.shape)  # Should output (44, 1, 14)

# Define the LSTM model
model = Sequential()
model.add(LSTM(4, stateful=True, batch_input_shape=(1, X_train.shape[1], X_train.shape[2])))
model.add(Dense(1))
model.compile(loss='mean_squared_error', optimizer='adam')

# Train the model
model.fit(X_train, y_train, epochs=100, batch_size=1, verbose=1, shuffle=False)

# Summarize the model
model.summary()

已尝试操作

  • 验证X_train形状为(44, 1, 14);
  • 尝试用input_shape替代batch_input_shape,但出现其他错误;
  • 确认TensorFlow与Keras版本兼容。

系统信息

  • Python版本:3.12
  • TensorFlow版本:2.17.0
  • Keras版本:3.4.1

解决方案

原因说明

Keras 3.x(TensorFlow 2.16及以上版本对应)已经移除了batch_input_shape参数,该参数不再被LSTM层识别,这是触发错误的核心原因。

两种修正方案

方案1:使用Input层显式指定batch_shape

将模型定义部分修改为:

# Define the LSTM model
model = Sequential()
# 用Input层指定批量形状:(batch_size, time_steps, features)
model.add(Input(batch_shape=(1, X_train.shape[1], X_train.shape[2])))
model.add(LSTM(4, stateful=True))
model.add(Dense(1))
model.compile(loss='mean_squared_error', optimizer='adam')

方案2:在LSTM层中同时指定input_shape和batch_size

如果不想额外添加Input层,可以直接在LSTM层中声明input_shape和batch_size参数:

# Define the LSTM model
model = Sequential()
# input_shape指定(time_steps, features),batch_size指定批量大小
model.add(LSTM(4, stateful=True, input_shape=(X_train.shape[1], X_train.shape[2]), batch_size=1))
model.add(Dense(1))
model.compile(loss='mean_squared_error', optimizer='adam')

关键注意事项

  1. 保持批量大小一致:训练时model.fit的batch_size必须和你指定的batch_size(或batch_shape中的第一个值)完全相同,你当前设置的batch_size=1符合要求;
  2. 禁用打乱数据:有状态LSTM依赖序列的连续性,必须设置shuffle=False,你的代码已经正确配置;
  3. 重置状态:如果需要在每个epoch结束后重置模型状态,可以在训练循环中手动调用model.reset_states(),例如:
for epoch in range(100):
    model.fit(X_train, y_train, epochs=1, batch_size=1, verbose=1, shuffle=False)
    model.reset_states()

内容的提问来源于stack exchange,提问作者AmaniAli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 06:00:54