You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

绘制线性回归最佳拟合线时y轴范围异常过小的解决方法

问题描述

我拥有x和y两个DataFrame数据,绘制散点图时结果表现正常(如图所示)。但将数据拟合到回归模型并绘制最佳拟合线后,拟合线的数值明显过高,导致y轴被挤压成密集的簇状(如图所示)。请问如何让y轴恢复正常范围?

以下是我的代码:

x = df["Year"]
y = df["Top speed"]

reg_prep = LinearRegression()

mod_reg = reg_prep.fit(x.to_numpy().reshape((-1,1)),y.to_numpy())
plt.scatter(x,y)

b0 = mod_reg.intercept_
b1 = mod_reg.coef_[0]
yfit = [b0 + b1 * xi for xi in x]
plt.plot(x,np.array(yfit).reshape(-1,1))
plt.show()
解决方法

问题核心是年份(x)数值过大,导致回归截距b0异常偏高,拟合线的数值范围被不合理拉高,挤压了原始数据的显示空间。可以通过两种方式解决:

  • 对年份做中心化处理
    将年份减去基准值(比如数据中的最小年份),缩小x的数值范围,让截距回归到有实际意义的区间,拟合线就会贴合原始数据的y轴范围。修改后的代码:

    import numpy as np
    from sklearn.linear_model import LinearRegression
    import matplotlib.pyplot as plt
    
    x = df["Year"]
    # 中心化处理,让x起始值接近0
    x_centered = x - x.min()
    y = df["Top speed"]
    
    reg_prep = LinearRegression()
    mod_reg = reg_prep.fit(x_centered.to_numpy().reshape((-1,1)), y.to_numpy())
    plt.scatter(x, y)
    
    b0 = mod_reg.intercept_
    b1 = mod_reg.coef_[0]
    yfit = [b0 + b1 * xi for xi in x_centered]
    plt.plot(x, np.array(yfit))
    plt.show()
    
  • 使用模型自带的predict方法生成拟合值
    无需手动计算截距和系数,直接调用sklearn模型的predict方法,既避免手动计算失误,结果也更可靠:

    import numpy as np
    from sklearn.linear_model import LinearRegression
    import matplotlib.pyplot as plt
    
    x = df["Year"].to_numpy().reshape((-1,1))
    y = df["Top speed"]
    
    reg_prep = LinearRegression()
    mod_reg = reg_prep.fit(x, y)
    plt.scatter(x, y)
    
    # 直接生成拟合值
    yfit = mod_reg.predict(x)
    plt.plot(x, yfit)
    plt.show()
    

另外注意检查代码是否导入了numpy和matplotlib.pyplot,原代码使用了np.array但未显示导入,这可能引发潜在错误。

内容的提问来源于stack exchange,提问作者km123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 14:30:40