You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python通过最佳拟合线预测现有DataFrame中变量的值?

使用线性拟合预测趋势并绘图

步骤分解与代码实现

首先导入所需库:

import pandas as pd
from scipy.stats import linregress
import matplotlib.pyplot as plt

1. 准备数据

替换成你自己的真实DataFrame即可,这里先给出示例结构:

# 构造截止到2013年的示例数据
data = {
    'Year': [2000, 2003, 2006, 2009, 2012, 2013],
    'Income': [35000, 38000, 42000, 45000, 48000, 49000]
}
df = pd.DataFrame(data)

2. 计算线性回归参数

用linregress()算出拟合线的核心参数:

# 提取变量
x = df['Year']
y = df['Income']

# 执行线性回归计算
reg_result = linregress(x, y)
slope = reg_result.slope    # 拟合线斜率
intercept = reg_result.intercept  # 拟合线截距

拟合公式为:预测收入 = 斜率 * 年份 + 截距

3. 生成预测年份与对应收入

创建包含原始年份和2014-2050年的序列,代入公式计算预测值:

# 生成从最早年份到2050的完整年份序列
prediction_years = pd.Series(range(x.min(), 2051))

# 计算每个年份的预测收入
predicted_income = slope * prediction_years + intercept

4. 绘制原始数据与预测趋势

用matplotlib把原始数据、拟合线和预测部分可视化:

plt.figure(figsize=(10, 6))

# 绘制原始数据散点
plt.scatter(x, y, color='darkblue', label='历史收入数据')

# 绘制完整的拟合预测线
plt.plot(prediction_years, predicted_income, color='crimson', linestyle='--', label='拟合预测趋势')

# 图表基础设置
plt.xlabel('年份')
plt.ylabel('收入')
plt.title('收入趋势预测(至2050年)')
plt.legend()
plt.grid(alpha=0.3)

plt.show()

额外补充

如果需要评估拟合效果,可以用linregress()返回的其他参数:

print(f"拟合线公式: Income = {slope:.2f} * Year + {intercept:.2f}")
print(f"相关系数(拟合度): {reg_result.rvalue:.2f}")

相关系数越接近1/-1,说明线性拟合效果越好。

内容的提问来源于stack exchange,提问作者Deefort

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:01:05