LSTM多变量预测新冠病例全为0问题排查求助
问题:多变量LSTM预测新冠新增病例全为0的排查与解决
基于OWID数据集构建LSTM模型预测新冠新增病例数,使用包含date、new_cases等6列的多变量序列时预测结果全为0,但单变量(仅total_cases)时无此问题。代码如下:
url = "https://covid.ourworldindata.org/data/owid-covid-data.csv" df = pd.read_csv(url) # Filter the data for a specific location, e.g., 'United States' location = 'United States' df_location = df[df['location'] == location] # Select relevant columns and set the date column as the index selected_columns = ['date', 'new_cases', 'new_deaths', 'total_cases', 'total_deaths', 'reproduction_rate'] df_location = df_location[selected_columns] df_location['date'] = pd.to_datetime(df_location['date']) df_location.set_index('date', inplace=True) # Handle missing values by filling them with the mean of the column df_location.fillna(df_location.mean(), inplace=True) # Function to create a multivariate dataset for LSTM def create_multivariate_dataset(data, time_step=1): X, Y = [], [] for i in range(len(data) - time_step - 1): X.append(data[i:(i + time_step)]) Y.append(data[i + time_step, 0]) # Predicting 'new_cases' return np.array(X), np.array(Y) # Convert the data to numpy array and scale it scaler = MinMaxScaler(feature_range=(0, 1)) scaled_data = scaler.fit_transform(df_location) # Create the dataset with a specified time step, e.g., 60 days time_step = 60 X, y = create_multivariate_dataset(scaled_data, time_step) # Reshape the input to be [samples, time steps, features] for LSTM X = X.reshape(X.shape[0], X.shape[1], X.shape[2]) # Split the data into training and testing sets X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Define the LSTM model model = Sequential() model.add(LSTM(100, return_sequences=True, input_shape=(time_step, X.shape[2]))) # Adjusted input_shape model.add(LSTM(100, return_sequences=False)) model.add(Dense(50, activation='relu')) model.add(Dense(1, activation='relu')) # Compile the model model.compile(optimizer='adam', loss='mean_squared_error') # Train the model history = model.fit(X_train, y_train, validation_data=(X_test, y_test), epochs=100, batch_size=32, verbose=1) # Make predictions on the test set test_predict = model.predict(X_test) # Inverse transform the predictions to get actual values test_predict_full = np.concatenate((test_predict, np.zeros((test_predict.shape[0], scaled_data.shape[1] - 1))), axis=1) test_predict = scaler.inverse_transform(test_predict_full)[:,0]
核心问题与解决办法
1. 输出层激活函数选择错误
你的输出层使用了relu激活函数,而new_cases经过MinMaxScaler缩放后范围是[0,1]。relu的特性是当输入小于0时输出0,若模型学习过程中权重导致输出值为负,就会被直接截断为0;即使输出接近0,也容易被压到0值区间。回归任务中,输出层不需要激活函数限制范围,linear才是正确选择。
修改代码:
model.add(Dense(1, activation='linear')) # 替换原有的relu激活函数
2. 时间序列数据拆分方式错误
你用train_test_split拆分数据会打乱时序顺序,而LSTM是依赖时序依赖关系的模型,打乱后模型无法学习到正确的时间趋势,直接导致预测失效。必须按时间先后顺序拆分训练集和测试集。
修改代码:
# 替换原有的train_test_split代码 train_size = int(len(X) * 0.8) X_train, X_test = X[:train_size], X[train_size:] y_train, y_test = y[:train_size], y[train_size:]
3. 逆缩放逻辑错误
你在逆缩放时将其他特征补为0,但MinMaxScaler是基于所有特征的整体分布进行缩放的,补0会引入错误的分布信息,导致逆变换结果失真。正确的做法是分离特征和目标变量的缩放,单独处理目标值的逆变换。
推荐修改方案:
# 拆分特征与目标变量 features = df_location.drop('new_cases', axis=1).values target = df_location['new_cases'].values.reshape(-1, 1) # 分别初始化缩放器 scaler_features = MinMaxScaler(feature_range=(0, 1)) scaler_target = MinMaxScaler(feature_range=(0, 1)) # 分别缩放特征和目标 scaled_features = scaler_features.fit_transform(features) scaled_target = scaler_target.fit_transform(target) # 合并为多变量序列 scaled_data = np.concatenate([scaled_target, scaled_features], axis=1) # 后续预测后,直接用目标变量的缩放器逆变换 test_predict = scaler_target.inverse_transform(test_predict)
4. 缺失值处理方式不合理
用列均值填充时序数据的缺失值会破坏数据的时间趋势,比如reproduction_rate这类具有时序特性的变量,均值填充会抹平其波动,影响模型学习。更适合用向前填充或线性插值来保留时序特征。
修改代码:
# 向前填充(保留最近的有效数据) df_location.fillna(method='ffill', inplace=True) # 或者用线性插值(更适合连续趋势的变量) # df_location.interpolate(method='linear', inplace=True)
内容的提问来源于stack exchange,提问作者Muhamed Khaled
相关产品推荐
相关产品推荐

