如何在Python Pandas中计算斜率、截距并生成y=mx+c列?
在Pandas中计算斜率(m)、截距(c)并生成
y=mx+c列 完全可以实现,分两种场景给你具体方案:
场景1:已知m和c为固定常数
如果m、c是预先确定的固定值,直接通过Pandas的列运算就能新增y列:
import pandas as pd from datetime import date import numpy as np # 生成你提供的示例数据 date_range = pd.date_range(date(2021,11,7), date.today()) value = np.random.rand(len(date_range)) historical = pd.DataFrame({'date': date_range, 'Sales' : value}) # 设定固定的m和c(示例值,可根据需求修改) m = 1.5 c = 0.3 # 方案A:以Sales作为x变量计算y historical['y'] = m * historical['Sales'] + c # 方案B:以日期作为x变量(需先转成可计算的数值,比如距离起始日的天数) historical['x_date'] = (historical['date'] - historical['date'].min()).dt.days historical['y_from_date'] = m * historical['x_date'] + c
场景2:通过现有数据拟合得到m和c
如果需要基于数据(比如日期和Sales的线性关系)拟合出回归方程的m和c,可用两种常用方法:
方法1:用numpy.polyfit快速拟合
# 先把日期转为数值型x historical['x_date'] = (historical['date'] - historical['date'].min()).dt.days # 拟合一次线性方程,返回结果就是[m, c] m, c = np.polyfit(historical['x_date'], historical['Sales'], deg=1) # 生成拟合后的y列 historical['y_fit'] = m * historical['x_date'] + c
方法2:用sklearn线性回归(适合复杂场景扩展)
from sklearn.linear_model import LinearRegression # 准备特征矩阵(sklearn要求输入为二维数组) X = historical['x_date'].values.reshape(-1, 1) y_target = historical['Sales'] # 训练模型 model = LinearRegression() model.fit(X, y_target) # 提取斜率和截距 m = model.coef_[0] c = model.intercept_ # 生成拟合列 historical['y_fit'] = model.predict(X)
示例输出效果
执行代码后,你的DataFrame会新增对应列,类似:
date Sales x_date y_fit 0 2021-11-07 0.210038 0 0.207341 1 2021-11-08 0.918222 1 0.208423 2 2021-11-09 0.202677 2 0.209505 3 2021-11-10 0.620185 3 0.210587 4 2021-11-11 0.299857 4 0.211669
内容的提问来源于stack exchange,提问作者navee pp
相关产品推荐
相关产品推荐

