You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas DataFrame最后6列计算slope、rsq、p_value列?

问题描述

给定如下数据:

dd = {'d_2023-09': {0: 23.0, 1: 14.0, 2: 15.0, 3: 10.0}, 'd_2023-10': {0: 20.0, 1: 8.0, 2: 7.0, 3: 12.0}, 'd_2023-11': {0: 14.0, 1: 9.0, 2: 11.0, 3: 11.0}, 'd_2023-12': {0: 19.0, 1: 17.0, 2: 13.0, 3: 7.0}, 'd_2024-01': {0: 18.0,
1: 11.0, 2: 19.0, 3: 10.0}, 'd_2024-02': {0: 13.0, 1: 4.0, 2: 6.0, 3: 13.0}, 'd_2024-03': {0: 21.0, 1: 10.0, 2: 11.0, 3: 6.0}}
df = pd.DataFrame(dd)

需要新增3列slope、rsq和pvalue,分别对应每行数据的最后6列与[1,2,3,4,5,6]做线性回归后的斜率、决定系数和p值。

尝试了以下代码但未成功:

df['slope']= np.polyfit(df[[i for i in dd.columns][-6:]],[1,2,3,4,5,6],1)[0]
解决方案

你的代码失效原因是np.polyfit默认按列处理数据,且无法直接返回rsq和pvalue。推荐用scipy.stats.linregress实现逐行回归并获取所有需要的统计量,具体步骤如下:

  1. 导入所需依赖库
import pandas as pd
import numpy as np
from scipy.stats import linregress
  1. 定义处理每行数据的函数
def get_reg_stats(row):
    # 提取当前行的最后6列数据作为因变量y
    y = row[-6:]
    # 定义自变量x
    x = np.array([1,2,3,4,5,6])
    # 执行线性回归
    reg_result = linregress(x, y)
    # 计算决定系数rsq(相关系数的平方)
    r_squared = reg_result.rvalue ** 2
    # 返回需要的三个统计量
    return pd.Series([reg_result.slope, r_squared, reg_result.pvalue], 
                     index=['slope', 'rsq', 'pvalue'])
  1. 逐行应用函数并合并结果
# 按行应用回归函数
reg_stats_df = df.apply(get_reg_stats, axis=1)
# 将统计量列合并到原DataFrame
df = pd.concat([df, reg_stats_df], axis=1)

关键说明

  • linregress(x, y)会返回包含斜率、截距、相关系数、p值、标准误的结果对象,能一次性满足需求
  • axis=1参数确保函数对每行数据单独处理,而非默认的按列处理
  • 决定系数rsq通过相关系数的平方计算,用于衡量回归模型的拟合程度

内容的提问来源于stack exchange,提问作者frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 20:35:19