如何将预测值添加到Pandas DataFrame中与原时间序列一同绘制
你遇到的报错是因为df['Interest_Rate']是pandas的Series对象,直接append普通Python列表会触发类型不匹配错误,而且pandas后续版本已经废弃了Series.append方法,更推荐先构造结构化的新数据框,再和原数据集合并。
完整修正代码
import pandas as pd from sklearn import linear_model import statsmodels.api as sm Stock_Market = {'Year': [2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016], 'Month': [12, 11,10,9,8,7,6,5,4,3,2,1,12,11,10,9,8,7,6,5,4,3,2,1], 'Interest_Rate': [2.75,2.5,2.5,2.5,2.5,2.5,2.5,2.25,2.25,2.25,2,2,2,1.75,1.75,1.75,1.75,1.75,1.75,1.75,1.75,1.75,1.75,1.75], 'Unemployment_Rate': [5.3,5.3,5.3,5.3,5.4,5.6,5.5,5.5,5.5,5.6,5.7,5.9,6,5.9,5.8,6.1,6.2,6.1,6.1,6.1,5.9,6.2,6.2,6.1], 'Stock_Index_Price': [1464,1394,1357,1293,1256,1254,1234,1195,1159,1167,1130,1075,1047,965,943,958,971,949,884,866,876,822,704,719] } df = pd.DataFrame(Stock_Market,columns=['Year','Month','Interest_Rate','Unemployment_Rate','Stock_Index_Price']) X = df[['Interest_Rate','Unemployment_Rate']] Y = df['Stock_Index_Price'] # 训练模型 regr = linear_model.LinearRegression() regr.fit(X, Y) print('Intercept: \n', regr.intercept_) print('Coefficients: \n', regr.coef_) # ---------------- 新增部分:构造预测数据集并合并 ---------------- # 新的特征数据 New_Interest_Rate = [2.75, 3, 4, 1, 2] New_Unemployment_Rate = [5.3, 4, 3, 2, 1] # 批量预测5组结果,避免循环 new_X = pd.DataFrame({'Interest_Rate': New_Interest_Rate, 'Unemployment_Rate': New_Unemployment_Rate}) new_predict = regr.predict(new_X) # 构造新数据的完整DataFrame,和原数据集列对齐 # 这里Year和Month按原数据时间顺推,原数据最后一条是2016年1月,新数据依次为2016年2-6月,你也可以根据自己的需求修改 new_df = pd.DataFrame({ 'Year': [2016]*5, 'Month': [2,3,4,5,6], 'Interest_Rate': New_Interest_Rate, 'Unemployment_Rate': New_Unemployment_Rate, 'Stock_Index_Price': new_predict }) # 合并原数据集和新预测数据集,ignore_index重置行号 full_df = pd.concat([df, new_df], ignore_index=True) # 打印合并后的最后10行验证结果 print(full_df.tail(10)) # ---------------- 以下原有statsmodels部分按需保留 ---------------- X_sm = sm.add_constant(X) model = sm.OLS(Y, X_sm).fit() predictions = model.predict(X_sm) print_model = model.summary() print(print_model)
关键修改说明
- 直接批量传入特征矩阵给
regr.predict,不需要循环逐行预测,效率更高 - 新数据先构造成和原数据集列完全一致的DataFrame,再用
pd.concat合并,不会触发类型错误 - 合并时设置
ignore_index=True可以自动重置整个数据集的行索引,避免索引重复
内容的提问来源于stack exchange,提问作者user2543
相关产品推荐
相关产品推荐

