基于RNN(LSTM)的时间序列数据集增强异常问题求助
时间序列合成数据后期趋同问题的解决思路
问题背景
使用LSTM模型增强IoT传感器生成的1小时间隔时间序列数据,目标是将20/67/280天的训练数据集扩展X天合成数据。但生成的合成数据后期全部变为相同值,即使将训练轮次从100增至200仍未解决。原始数据无缺失,呈线性特征。
数据样本
原始数据
timestamp,value 2013-07-04 00:00:00,69.88083514 2013-07-04 01:00:00,71.22022706 2013-07-04 02:00:00,70.87780496 2013-07-04 03:00:00,68.95939994 2013-07-04 04:00:00,69.28355102 2013-07-04 05:00:00,70.06096581 2013-07-04 06:00:00,69.27976479 2013-07-04 07:00:00,69.36960846 ....
合成数据异常表现
.... 2014-03-02 02:00:00,65.00628144 2014-03-02 03:00:00,65.10276279 2014-03-03 09:00:00,64.73752596 .... 2014-03-03 09:00:00,64.90076446533203 2014-03-03 10:00:00,65.06415557861328 2014-03-03 11:00:00,65.22787475585938 2014-03-03 12:00:00,65.39212799072266 2014-03-03 13:00:00,65.5571060180664 2014-03-03 14:00:00,65.7229995727539 2014-03-03 15:00:00,65.8899917602539 .... 2014-03-09 09:00:00,85.6402587890625 2014-03-09 10:00:00,85.64026641845703 2014-03-09 11:00:00,85.64026641845703 2014-03-09 12:00:00,85.64026641845703 2014-03-09 13:00:00,85.64026641845703 2014-03-09 14:00:00,85.64026641845703 2014-03-09 15:00:00,85.64026641845703 2014-03-09 16:00:00,85.64026641845703
现有实现代码
import pandas as pd import numpy as np from keras.layers import LSTM, Dense from keras.models import Sequential def augment_timeseries(df, num_days): # Create a copy of the timestamp column timestamp_col = df['timestamp'].copy() # Convert timestamp column to datetime format df['timestamp'] = pd.to_datetime(df['timestamp']) # Set timestamp column as the index df.set_index('timestamp', inplace=True) # Get the frequency of the timestamps in the original dataset freq = 'H' # Create a range of timestamps with the same frequency as the original dataset end_time = df.index[-1] + pd.DateOffset(days=num_days) timestamps = pd.date_range(start=df.index[-1], end=end_time, freq=freq) # Reshape data for input into RNN model X_train = df.values.reshape((-1, 1, 1)) # Create the RNN model model = Sequential() model.add(LSTM(50, input_shape=(1, 1), return_sequences=True)) model.add(Dense(1)) model.compile(loss='mean_squared_error', optimizer='adam') # Fit the model on the data model.fit(X_train, X_train, epochs=100, batch_size=1, verbose=1) # Use the model to make predictions for new timestamps for timestamp in timestamps: X_test = np.array([[df.values[-1]]]) prediction = model.predict(X_test) df.loc[len(df)] = [prediction[0][0][0]] timestamp_df = pd.DataFrame([timestamp]) timestamp_col = pd.concat([timestamp_col, timestamp_df], ignore_index=True) model.reset_states() # Add the timestamp column back to the original dataframe df.insert(0, 'timestamp', timestamp_col.values) return df
解决思路与技术修正
- 重构训练数据输入逻辑:当前代码仅用单个值预测下一个值,且每次预测后重置模型状态,完全浪费了LSTM的序列记忆能力。应使用滑动窗口构建训练样本,比如用过去24小时(对应日周期)的数据作为输入,预测下一小时的值,让模型学习序列的趋势与周期特征。示例:将
X_train构造成(样本数, 时间步长, 特征数)的格式,时间步长设为24。 - 调整模型结构:当前LSTM设置
return_sequences=True但仅接Dense层,输出为序列而非单值。单步预测场景下,应将return_sequences设为False,或在最后一层LSTM后接Dense输出单值。同时可加入Dropout层(如model.add(Dropout(0.2)))防止过拟合。 - 添加数据归一化:LSTM对数据尺度敏感,建议用
MinMaxScaler将数据缩放到[0,1]区间,预测后再反缩放还原,避免数值漂移导致的收敛问题。 - 优化预测流程:取消每次预测后的
model.reset_states()调用,让模型保持对生成序列的记忆;或采用多步预测方式,一次生成多个时间步的结果,而非单步迭代。 - 调整训练策略:降低Adam优化器的学习率(如
learning_rate=0.0001),增大batch_size提升训练稳定性;加入验证集监控过拟合,避免盲目增加训练轮次。 - 适配线性特征:原始数据呈线性,可尝试先对数据做差分处理消除趋势,训练生成差分序列后再还原趋势;或在模型中结合线性层,提升对线性特征的拟合能力。
内容的提问来源于stack exchange,提问作者UberDataGeek
相关产品推荐
相关产品推荐

