如何在scikit-learn中为长时序数据构建回归预测模型?
修正后的需求预测回归模型方案
原代码核心问题
- 仅提取单一行数据(
data = df.iloc[0]),样本量为1,模型无法学习任何规律 - 先执行
model.fit(X_train, y_train)再拆分数据,此时X_train和y_train未定义,会直接报错 - 目标变量仅选取1980年的需求,无法捕捉时间维度的需求变化规律
- 宽格式数据结构不适合时间序列回归建模
完整修正代码
import pandas as pd from sklearn.ensemble import RandomForestRegressor from sklearn.model_selection import train_test_split from sklearn.metrics import r2_score # ---------------------- 1. 数据预处理:宽格式转长格式 ---------------------- # 假设原始df包含列:Country, ISO, 1980, 1981,...,2021, 1980_dem, 1981_dem,...,2021_dem year_cols = [str(year) for year in range(1980, 2022)] dem_cols = [f"{year}_dem" for year in range(1980, 2022)] # 重塑人口数据 population_df = df.melt(id_vars=['Country', 'ISO'], value_vars=year_cols, var_name='Year', value_name='Population') population_df['Year'] = population_df['Year'].astype(int) # 重塑需求数据 demand_df = df.melt(id_vars=['Country', 'ISO'], value_vars=dem_cols, var_name='Year_dem', value_name='Demand') demand_df['Year'] = demand_df['Year_dem'].str.extract('(\d+)').astype(int) demand_df = demand_df.drop('Year_dem', axis=1) # 合并人口与需求数据 processed_data = pd.merge(population_df, demand_df, on=['Country', 'ISO', 'Year'], how='inner') # 处理缺失值 processed_data = processed_data.dropna(subset=['Population', 'Demand']) # ---------------------- 2. 特征与目标变量定义 ---------------------- # 对国家进行数值编码 processed_data['Country_code'] = processed_data['Country'].astype('category').cat.codes X = processed_data[['Country_code', 'Year', 'Population']] y = processed_data['Demand'] # ---------------------- 3. 拆分数据并训练模型 ---------------------- X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=44) model = RandomForestRegressor(n_estimators=50, max_features="auto", random_state=44) model.fit(X_train, y_train) # ---------------------- 4. 模型评估 ---------------------- y_pred = model.predict(X_test) r2 = r2_score(y_test, y_pred) print(f"模型R²分数: {round(r2, 2)}") # ---------------------- 5. 未来需求预测 ---------------------- # 示例:构造某国家2022-2030年的人口预测数据 future_data = pd.DataFrame({ 'Country_code': [processed_data[processed_data['Country'] == 'China']['Country_code'].iloc[0]]*9, 'Year': range(2022, 2031), 'Population': [1412000000, 1410000000, 1408000000, 1405000000, 1402000000, 1399000000, 1396000000, 1393000000, 1390000000] }) # 预测未来需求 future_demand = model.predict(future_data) future_data['Predicted_Demand'] = future_demand print("\n未来需求预测结果:") print(future_data[['Year', 'Population', 'Predicted_Demand']])
关键说明
- 数据重塑:通过
melt将宽格式转为长格式,让模型能学习年份、人口与需求的关联 - 特征编码:将国家名称转为数值编码,适配树模型对特征类型的要求
- 评估指标:回归任务用R²分数评估(范围0-1,越接近1拟合效果越好),避免误用分类任务的accuracy指标
- 未来预测:需提前准备未来年份的人口数据(可通过专业人口预测数据源获取),构造与训练数据一致的特征格式即可生成需求预测
内容的提问来源于stack exchange,提问作者PiotrK
相关产品推荐
相关产品推荐

