You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Walk-Forward优化与多元线性回归:R²分数定义异常问题求助

电力消耗预测建模问题排查

问题背景

我尝试用gdp、pop、exp、sp_500、temp这5个特征预测电力消耗(ec),采用Walk-Forward验证策略:初始用350条月度数据训练,预测接下来6个月的数据;之后每次迭代新增1条训练数据,再预测后续6个月。

自定义Walk-Forward拆分代码

# initial: initial train length 
# horizon: forecast horizon (test set length). Default = 1
# period:  length of train data to add each iteration 

class expanding_window(object):
    
    def __init__(self,initial=1, horizon=1, period=1):
        self.initial = initial
        self.horizon = horizon
        self.period = period
        
    def split(self, data):
        self.data = data
        self.counter = 0 
        
        data_length = data.shape[0]
        data_index = list(np.arange(data_length))
        
        output_train = []
        output_test = []
        
        # append initial
        output_train.append(list(np.arange(self.initial)))
        progress = [x for x in data_index if x not in list(np.arange(self.initial))]
        output_test.append([x for x in data_index if x not in output_train[self.counter]][:self.horizon])
        
        while len(progress) !=0:
            temp = progress[:self.period]
            to_add = output_train[self.counter] + temp
            
            # update the train index
            output_train.append(to_add)
            
            # increment counter
            self.counter +=1
            
            # update test index
            to_add_test = [x for x in data_index if x not in output_train[self.counter]][:self.horizon]
            output_test.append(to_add_test)
            
            # update progress
            progress = [x for x in data_index if x not in output_train[self.counter]]
        
        # clip last element of output_train and output_test
        output_train = output_train[:-1]
        output_test = output_test[:-1]
        
        # mimic sklearn output
        index_output = [(train,test) for train,test in zip(output_train, output_test)]
        
        return index_output

# Separate the predictors and label
X = data[data.columns[~data.columns.isin(
    ["gdp", "pop", "exp", "sp_500", "temp"])]]
y = data["ec"]

tscv = expanding_window(initial=350, horizon = 6, period = 1)
for train_index, test_index in tscv.split(X):
    print("Train:",train_index)
    print("Test :",test_index)

for i, (train_index, test_index) in enumerate(tscv.split(data)):
    print("Split",i)
    print("Train:",data[["gdp","pop","exp","sp_500","temp"]].iloc[train_index])

多元线性回归(MLR)建模代码

# Multiple Linear Regression
from sklearn.linear_model import LinearRegression

ols = LinearRegression()
oos_score_list = []

print("Split  In-sample R^2  Out-of-Sample R^2")
print("-"*40)

# Loop through the splits. Run a Linear Regression for each split.
for i, (train_index, test_index) in enumerate(tscv.split(data)):
    X_train = data[["gdp", "pop", "exp", "sp_500", "temp"]].iloc[train_index]
    y_train = data[["ec"]].iloc[train_index]
    X_test = data[["gdp", "pop", "exp", "sp_500", "temp"]].iloc[test_index]
    y_test = data[["ec"]].iloc[test_index]
    ols.fit(X_train,y_train)
    oos_score = ols.score(X_test,y_test)
    print(i,
          " "*4,
          round(ols.score(X_train,y_train),2),
          " "*10, 
          round(oos_score,2))
    oos_score_list.append(oos_score)

oos_score_list = oos_score_list[ : -1]
print("-"*40)
print("Average out-of-sample score:",round(np.mean(oos_score_list),2))

遇到的问题

  • 运行MLR代码时收到警告:UndefinedMetricWarning: R^2 score is not well-defined with less than two samples
  • 样本外平均R²得分低于0(-0.52)

问题排查方向

1. Walk-Forward拆分逻辑错误

  • 特征与目标变量拆分完全颠倒:原代码中X = data[data.columns[~data.columns.isin(["gdp", "pop", "exp", "sp_500", "temp"])]]把需要的预测特征全部排除,正确写法应为X = data[["gdp", "pop", "exp", "sp_500", "temp"]],保留预测特征,y = data["ec"]作为目标变量。
  • 测试集样本数不足:最后几次迭代中,剩余数据可能无法凑够6个测试样本,导致test_index样本数<2,触发R²计算的警告。需在split方法中添加判断:仅当剩余数据≥horizon时才生成测试集,否则终止迭代。
  • 拆分数据集不统一:代码中分别用X和data调用tscv.split(),可能因数据集长度不同导致索引混乱,需统一用同一个数据集(如data)执行拆分。

2. 模型适配性问题

  • 未考虑时间序列特性:电力消耗属于时间序列数据,存在趋势、季节性、自相关性等特征,多元线性回归未捕捉这些规律,导致样本外表现极差。建议先分析数据的时间特性,加入月份/年份等时间特征,或改用ARIMA、Prophet等时间序列模型。
  • 特征与目标相关性极低:若5个特征与ec的线性相关性弱,线性回归无法有效预测。可通过data.corr()生成相关系数矩阵,验证特征与目标变量的线性关联程度。

3. 代码执行细节问题

  • 无效得分计算:当测试集样本数<2时,R²无法有效计算,需在循环中判断len(test_index)>=2,再执行得分计算,避免无效值拉低平均得分。

内容的提问来源于stack exchange,提问作者Freyana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 18:48:24