Walk-Forward优化与多元线性回归:R²分数定义异常问题求助
电力消耗预测建模问题排查
问题背景
我尝试用gdp、pop、exp、sp_500、temp这5个特征预测电力消耗(ec),采用Walk-Forward验证策略:初始用350条月度数据训练,预测接下来6个月的数据;之后每次迭代新增1条训练数据,再预测后续6个月。
自定义Walk-Forward拆分代码
# initial: initial train length # horizon: forecast horizon (test set length). Default = 1 # period: length of train data to add each iteration class expanding_window(object): def __init__(self,initial=1, horizon=1, period=1): self.initial = initial self.horizon = horizon self.period = period def split(self, data): self.data = data self.counter = 0 data_length = data.shape[0] data_index = list(np.arange(data_length)) output_train = [] output_test = [] # append initial output_train.append(list(np.arange(self.initial))) progress = [x for x in data_index if x not in list(np.arange(self.initial))] output_test.append([x for x in data_index if x not in output_train[self.counter]][:self.horizon]) while len(progress) !=0: temp = progress[:self.period] to_add = output_train[self.counter] + temp # update the train index output_train.append(to_add) # increment counter self.counter +=1 # update test index to_add_test = [x for x in data_index if x not in output_train[self.counter]][:self.horizon] output_test.append(to_add_test) # update progress progress = [x for x in data_index if x not in output_train[self.counter]] # clip last element of output_train and output_test output_train = output_train[:-1] output_test = output_test[:-1] # mimic sklearn output index_output = [(train,test) for train,test in zip(output_train, output_test)] return index_output # Separate the predictors and label X = data[data.columns[~data.columns.isin( ["gdp", "pop", "exp", "sp_500", "temp"])]] y = data["ec"] tscv = expanding_window(initial=350, horizon = 6, period = 1) for train_index, test_index in tscv.split(X): print("Train:",train_index) print("Test :",test_index) for i, (train_index, test_index) in enumerate(tscv.split(data)): print("Split",i) print("Train:",data[["gdp","pop","exp","sp_500","temp"]].iloc[train_index])
多元线性回归(MLR)建模代码
# Multiple Linear Regression from sklearn.linear_model import LinearRegression ols = LinearRegression() oos_score_list = [] print("Split In-sample R^2 Out-of-Sample R^2") print("-"*40) # Loop through the splits. Run a Linear Regression for each split. for i, (train_index, test_index) in enumerate(tscv.split(data)): X_train = data[["gdp", "pop", "exp", "sp_500", "temp"]].iloc[train_index] y_train = data[["ec"]].iloc[train_index] X_test = data[["gdp", "pop", "exp", "sp_500", "temp"]].iloc[test_index] y_test = data[["ec"]].iloc[test_index] ols.fit(X_train,y_train) oos_score = ols.score(X_test,y_test) print(i, " "*4, round(ols.score(X_train,y_train),2), " "*10, round(oos_score,2)) oos_score_list.append(oos_score) oos_score_list = oos_score_list[ : -1] print("-"*40) print("Average out-of-sample score:",round(np.mean(oos_score_list),2))
遇到的问题
- 运行MLR代码时收到警告:
UndefinedMetricWarning: R^2 score is not well-defined with less than two samples - 样本外平均R²得分低于0(-0.52)
问题排查方向
1. Walk-Forward拆分逻辑错误
- 特征与目标变量拆分完全颠倒:原代码中
X = data[data.columns[~data.columns.isin(["gdp", "pop", "exp", "sp_500", "temp"])]]把需要的预测特征全部排除,正确写法应为X = data[["gdp", "pop", "exp", "sp_500", "temp"]],保留预测特征,y = data["ec"]作为目标变量。 - 测试集样本数不足:最后几次迭代中,剩余数据可能无法凑够6个测试样本,导致
test_index样本数<2,触发R²计算的警告。需在split方法中添加判断:仅当剩余数据≥horizon时才生成测试集,否则终止迭代。 - 拆分数据集不统一:代码中分别用
X和data调用tscv.split(),可能因数据集长度不同导致索引混乱,需统一用同一个数据集(如data)执行拆分。
2. 模型适配性问题
- 未考虑时间序列特性:电力消耗属于时间序列数据,存在趋势、季节性、自相关性等特征,多元线性回归未捕捉这些规律,导致样本外表现极差。建议先分析数据的时间特性,加入月份/年份等时间特征,或改用ARIMA、Prophet等时间序列模型。
- 特征与目标相关性极低:若5个特征与ec的线性相关性弱,线性回归无法有效预测。可通过
data.corr()生成相关系数矩阵,验证特征与目标变量的线性关联程度。
3. 代码执行细节问题
- 无效得分计算:当测试集样本数<2时,R²无法有效计算,需在循环中判断
len(test_index)>=2,再执行得分计算,避免无效值拉低平均得分。
内容的提问来源于stack exchange,提问作者Freyana
相关产品推荐
相关产品推荐

