You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LinregressResult无groupby属性报错及分组线性回归斜率计算方案

问题与解决方案

问题描述

我有一个结构如下的pandas DataFrame:

imagefile   imagefile_mnemonic    a    b    pix    val
file1       image1                0    0    v1     55
file1       image1                0    1    v1     75
file1       image1                0    2    v1     95
file1       image1                0    3    v1     115
file1       image1                0    4    v1     135
file1       image1                0    5    v1     155
file1       image1                0    6    v1     175
...         ...                  ...  ...  ...    ...
file6       image6                23   11   v5     763
file6       image6                23   12   v5     787

需求:

  • 按imagefile_mnemonic、相同a值、相同pix类型分组
  • 以b列为x、val列为y计算线性回归斜率
  • 后续计算单个文件内所有a值对应斜率的平均值

尝试代码:

from scipy.stats import linregress
df_v1 = df[df['pix']=='v1']
slope, intercept, r_value, p_value, std_err = linregress(df_v1['b'], df_v1['val']).groupby(df_p['a'])

报错:

AttributeError: 'LinregressResult' object has no attribute 'groupby'

要求:在仍使用linregress的前提下解决问题。

解决方法

错误根源是操作顺序颠倒——linregress只能对一组数据计算回归结果,不能先调用它再执行groupby。正确逻辑是先按分组条件拆分数据,再对每个分组应用linregress提取斜率。

1. 按指定条件分组计算斜率

直接对原DataFrame按imagefile_mnemonic、a、pix分组,用apply对每个分组执行线性回归并提取斜率:

from scipy.stats import linregress
import pandas as pd

# 定义函数:输入分组数据,返回线性回归斜率
def get_slope(group):
    # 转换为数值型并过滤缺失值
    x = pd.to_numeric(group['b'], errors='coerce')
    y = pd.to_numeric(group['val'], errors='coerce')
    mask = ~x.isna() & ~y.isna()
    
    # 至少需要2个有效数据点才能计算回归
    if len(x[mask]) < 2:
        return None
    
    reg_result = linregress(x[mask], y[mask])
    return reg_result.slope

# 分组计算斜率
slope_df = df.groupby(['imagefile_mnemonic', 'a', 'pix']).apply(get_slope).reset_index(name='slope')

2. 计算单个文件的斜率平均值

基于slope_df,按imagefile_mnemonic分组求斜率均值:

file_avg_slope = slope_df.groupby('imagefile_mnemonic')['slope'].mean().reset_index(name='avg_slope')

补充说明

  • 加入了缺失值过滤和样本量判断,避免因数据异常导致计算报错
  • 如果需要保留回归的其他指标(如截距、r²值),可以修改get_slope函数,返回包含多指标的Series,再整理成结构化DataFrame

内容的提问来源于stack exchange,提问作者mycodeQ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 11:09:52