Python线性回归建模如何筛选1979-2020年的逐月数据子集
解决方案
首先你原有代码存在一处变量名笔误:读取csv时将数据存入了data1变量,后续构造DataFrame时错误调用了未定义的data,修正后再添加年份筛选逻辑即可实现需求,完整可运行代码如下:
import pandas as pd from sklearn import linear_model from sklearn.model_selection import train_test_split import statsmodels.api as sm import matplotlib.pyplot as plt from sklearn.linear_model import LinearRegression clf = LinearRegression() # 读取数据 data1 = pd.read_csv (r'C:\Users\User-PC\sample.csv') x1 = pd.DataFrame(data1,columns=['Year','Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov','Dec']) # 新增:筛选1979-2020年的数据 filtered_data = x1[(x1['Year'] >= 1979) & (x1['Year'] <= 2020)] # 绘图使用筛选后的数据 plt.plot(filtered_data['Year'], filtered_data['Jan'], color='green') plt.title('Model 1') plt.xlabel('Year') plt.ylabel('index') plt.show() # 线性回归示例:以年份为自变量,1月数据为因变量拟合 # 构造特征和标签 X = filtered_data[['Year']] y = filtered_data['Jan'] # 拟合模型 clf.fit(X, y) # 输出回归结果 print(f"回归系数:{clf.coef_[0]}") print(f"截距:{clf.intercept_}") print(f"模型R²得分:{clf.score(X, y)}")
改动说明
- 修正了原代码的变量名笔误,保证数据读取逻辑正常
- 用pandas布尔索引实现年份范围筛选,得到的
filtered_data就是1979-2020年的全量数据,你后续所有分析运算都可以基于这个数据集开展 - 补充了最基础的单变量线性回归调用示例,如果你需要用12个月的数据作为特征做其他回归任务,只需调整
X和y的取值即可
内容的提问来源于stack exchange,提问作者user2543
相关产品推荐
相关产品推荐

