You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python线性回归建模如何筛选1979-2020年的逐月数据子集

解决方案

首先你原有代码存在一处变量名笔误:读取csv时将数据存入了data1变量,后续构造DataFrame时错误调用了未定义的data,修正后再添加年份筛选逻辑即可实现需求,完整可运行代码如下:

import pandas as pd
from sklearn import linear_model
from sklearn.model_selection import train_test_split
import statsmodels.api as sm
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
clf = LinearRegression()

# 读取数据
data1 = pd.read_csv (r'C:\Users\User-PC\sample.csv') 
x1 = pd.DataFrame(data1,columns=['Year','Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov','Dec'])

# 新增:筛选1979-2020年的数据
filtered_data = x1[(x1['Year'] >= 1979) & (x1['Year'] <= 2020)]

# 绘图使用筛选后的数据
plt.plot(filtered_data['Year'], filtered_data['Jan'], color='green')
plt.title('Model 1')
plt.xlabel('Year')
plt.ylabel('index')
plt.show()

# 线性回归示例:以年份为自变量,1月数据为因变量拟合
# 构造特征和标签
X = filtered_data[['Year']]
y = filtered_data['Jan']
# 拟合模型
clf.fit(X, y)
# 输出回归结果
print(f"回归系数:{clf.coef_[0]}")
print(f"截距:{clf.intercept_}")
print(f"模型R²得分:{clf.score(X, y)}")

改动说明

  • 修正了原代码的变量名笔误,保证数据读取逻辑正常
  • 用pandas布尔索引实现年份范围筛选,得到的filtered_data就是1979-2020年的全量数据,你后续所有分析运算都可以基于这个数据集开展
  • 补充了最基础的单变量线性回归调用示例,如果你需要用12个月的数据作为特征做其他回归任务,只需调整X和y的取值即可

内容的提问来源于stack exchange,提问作者user2543

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 16:48:01