You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将按PT_ID分组后的GroupBy对象转换为可分析的DataFrame?

问题解决:将GroupBy对象转换为DataFrame

你的核心问题是代码中把df赋值成了groupby返回的GroupBy对象,而非DataFrame。以下是两种直接解决的方案:

方案一:基于原有步骤拼接分组数据

沿用你先全局按P降序排序、再分组的逻辑,只需将分组后的各组数据重新拼接成DataFrame:

修改后的完整代码:

import pandas as pd

def preProcessing(df):
    """
    The preprocessing done on the Dataset for efficient use.
    @returns- The new and improved dataset
    We're using the following steps- 
    Sort Data by P
    Group by PT_ID and convert back to DataFrame
    
    """
    # 按P值降序排序
    df = df.sort_values(by='P', ascending=False) 
    # 分组后拼接所有组为DataFrame,sort=False保留PT_ID在全局排序后的出现顺序
    df = pd.concat([group for _, group in df.groupby('PT_ID', sort=False)])
    return df

方案二:更高效的组内排序方式

如果不需要全局先按P排序,仅要求每个PT_ID组内数据按P降序排列,可以直接通过多键排序实现,避免分组操作:

修改后的代码:

import pandas as pd

def preProcessing(df):
    """
    The preprocessing done on the Dataset for efficient use.
    @returns- The new and improved dataset
    We're using the following steps- 
    Sort Data by PT_ID, then by P in descending order
    
    """
    # 先按PT_ID聚合(逻辑上),再按P值降序排序,直接得到目标DataFrame
    df = df.sort_values(by=['PT_ID', 'P'], ascending=[True, False])
    return df

两种方案都能返回符合要求的DataFrame,组内数据保持P降序,各组顺序可根据需求调整(方案一中的sort=False会保留全局排序后PT_ID的出现顺序,若不需要可删除该参数,默认按PT_ID升序排列)。

内容的提问来源于stack exchange,提问作者pasha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:45:39