OLS回归预测输出开头出现多余0,求格式调整方案
问题:OLS回归预测结果格式不符合预期
使用statsmodels进行OLS简单线性回归,输入学士占比预测互联网使用率时,输出格式不符合预期:
- 用
pd.DataFrame打印会显示多余的列名0 - 直接打印预测结果是数组形式
- 预期仅显示行索引
0与对应预测值
代码示例
import numpy as np import pandas as pd import statsmodels.api as sm # 加载数据集 internet = pd.read_csv('internetusage.csv') # 提取变量 y = internet['internet_usage'] x = internet['bachelors_degree'] # 添加常数项 x = sm.add_constant(x) # 训练模型 model = sm.OLS(y, x).fit() # 获取输入并预测 bach_percent = float(input()) prediction = model.predict([1,bach_percent]) # 当前打印方式 print(pd.DataFrame(prediction)) print('dtype:', internet['internet_usage'].dtypes)
当前输出
0 0 61.865155 dtype: float64
预期输出
0 61.865155 dtype: float64
解决方法
方法1:直接格式化输出索引与预测值
因为prediction是仅含单个元素的数组,直接提取元素并按预期格式打印:
print(f"0 {prediction[0]:.6f}") print('dtype:', internet['internet_usage'].dtypes)
方法2:转为Series而非DataFrame
Series是一维结构,不会生成额外列名,通过to_string(header=False)隐藏默认标题:
pred_series = pd.Series(prediction, index=[0]) print(pred_series.to_string(header=False)) print('dtype:', internet['internet_usage'].dtypes)
方法3:调整DataFrame打印参数
如果坚持使用DataFrame,用to_string(header=False)隐藏列名:
df = pd.DataFrame(prediction) print(df.to_string(header=False)) print('dtype:', internet['internet_usage'].dtypes)
内容的提问来源于stack exchange,提问作者Victor Montanez
相关产品推荐
相关产品推荐

