如何改进LSTM模型以提升时间序列温度预测精度?
温度时间序列LSTM模型优化方案
核心问题识别与修改建议
1. 任务类型匹配错误(最致命)
温度预测属于回归任务,但当前代码误用了二分类任务的损失函数和评价指标:
- 替换
binary_crossentropy为回归任务常用的mse(均方误差)或mae(平均绝对误差) - 替换
accuracy评价指标为mae、mse或mape(平均绝对百分比误差) - 将模型Checkpoint监控的
val_accuracy改为val_loss,确保保存最优的回归模型权重
2. LSTM输入格式错误
LSTM要求输入维度为[样本数, 时间步长, 特征数],但当前代码未构造时间序列滑动窗口,直接传入一维数据,完全不符合模型输入要求:
- 需编写函数将原始序列转换为滑动窗口格式,用过去N个时间步的温度数据预测下一个时间步的温度,示例逻辑:
def create_sequences(data, seq_length): X, y = [], [] for i in range(len(data) - seq_length): X.append(data[i:i+seq_length]) y.append(data[i+seq_length]) return np.array(X), np.array(y) # 构造训练/验证集序列(示例用过去10个时间步预测下一个) Xtra_seq, Ytra_seq = create_sequences(Xtra, seq_length=10) Xval_seq, Yval_seq = create_sequences(Xval, seq_length=10)
3. 重采样参数错误
注释标注要转为2小时数据,但代码中用了resample('48H')(48小时=2天),过度压缩数据会丢失时间序列关键细节,应改为resample('2H').mean()
4. 特征单一化问题
仅用温度自身作为输入,无法捕捉时间序列的周期性(季节、月份)和趋势特征:
- 从时间索引中提取衍生特征,示例代码:
df['month'] = df.index.month df['quarter'] = df.index.quarter input_columns = ['warm well temperature (°C)', 'month', 'quarter'] X = df[input_columns]
5. 模型与训练参数优化
- LSTM循环层激活函数错误:默认
tanh更适合循环网络,替换relu为tanh - 学习率过高:
Adam(learning_rate=0.01)易导致训练震荡,建议降低到0.001或0.0001 - 添加早停机制防止过拟合:
from tensorflow.keras.callbacks import EarlyStopping early_stop = EarlyStopping(monitor='val_loss', patience=20, restore_best_weights=True) # 训练时加入callbacks=[modelCheckpoint, early_stop]
- 调整模型容量:当前
H=150可能过大,可尝试减小到32/64,或添加Dropout层抑制过拟合:
x = LSTM(H, activation='tanh', return_sequences=False)(i) x = Dropout(0.2)(x) # 添加Dropout层
原代码
#### Link to Google Drive and load modules # generic modules import datetime, os import itertools import time import pickle # basic data science modules import numpy as np import pandas as pd import matplotlib.pyplot as plt %matplotlib inline import seaborn as sns # keras from tensorflow.keras.layers import Input, Dense, Dropout from tensorflow.keras.models import Model from tensorflow.keras.optimizers import Adam # sklearn helper functions from sklearn.preprocessing import MinMaxScaler from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score, f1_score, classification_report, confusion_matrix # Import LSTM-related modules from tensorflow.keras.layers import LSTM df = pd.read_csv('Master CSV v1.csv',index_col='Date/Time',parse_dates=True) # Convert 5-minute data to 48-hour data # Resample the dataframe to 2-hour frequency and take the mean value within each 2-hour interval df = df.resample('48H').mean() df = pd.DataFrame(df, index=df.index) #### inputs and outpts Y = df['warm well temperature (°C)'] input_columns = ['warm well temperature (°C)'] X = df[input_columns] # Split the dataset into training and validation sets Xtra, Xval, Ytra, Yval = train_test_split(X, Y, test_size=0.20, shuffle=False) # Convert DataFrame to NumPy array Xtra = Xtra.values Xval = Xval.values # Rescale input and output datasets Sx = MinMaxScaler(feature_range=(0, 1)) Sy = MinMaxScaler(feature_range=(0, 1)) Xtra = Sx.fit_transform(Xtra.reshape(Xtra.shape[0], -1)) Xval = Sx.transform(Xval.reshape(Xval.shape[0], -1)) Ytra = Sy.fit_transform(Ytra.values.reshape(-1, 1)) Yval = Sy.transform(Yval.values.reshape(-1, 1)) # function for LSTM model def LSTMmodel(L, D, H): i = Input(shape=(L, D), name='input_layer') x = LSTM(H, activation='relu', name='recurrent_layer', return_sequences=False)(i) x = Dense(1, name='output_layer')(x) model = Model(i, x, name='LSTM') return model # Create an instance of the LSTM model lstm_model = LSTMmodel(L=130, D=1, H=150) # modelCheckpoint from tensorflow.keras.callbacks import ModelCheckpoint modelCheckpoint = ModelCheckpoint('LSTM_model_weights v1.hdf5', save_best_only=True, monitor='val_accuracy',mode='auto', save_weights_only=True) # Compile and train the LSTM model using the fit function lstm_model.compile(loss='binary_crossentropy', optimizer=Adam(learning_rate=0.01), metrics=['accuracy']) r = lstm_model.fit(Xtra, Ytra, epochs=500, validation_data=(Xval, Yval), verbose=0, batch_size=64, callbacks=[modelCheckpoint]) # load best model weights lstm_model.load_weights('LSTM_model_weights v1.hdf5') # Function for plotting training history def plot_training_history(r, figsize=(10, 3)): f, axes = plt.subplots(1, 2, figsize=figsize) # Loss axes[0].plot(r.history['loss'], label='Training Loss') axes[0].plot(r.history['val_loss'], label='Validation Loss') axes[0].set_title('Loss Trajectories') axes[0].set_xlabel('Epochs') axes[0].set_ylabel('Loss') axes[0].legend() # Accuracy axes[1].plot(r.history['accuracy'], label='Training Accuracy') axes[1].plot(r.history['val_accuracy'], label='Validation Accuracy') axes[1].set_title('Accuracy Trajectories') axes[1].set_xlabel('Epochs') axes[1].set_ylabel('Accuracy') axes[1].legend() # Adjust plot plt.tight_layout() plt.show() return f, axes # Call the modified function plot_training_history(r, figsize=(10, 3))
内容的提问来源于stack exchange,提问作者amir
相关产品推荐
相关产品推荐

