You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何展开groupby分组结果并与原DataFrame保持索引/顺序一致?

分组后结果与原DataFrame索引/顺序对齐问题

给定如下DataFrame和分组内执行的函数:

import pandas as pd

df = pd.DataFrame({"id": [1, 2, 1, 2], "amount": [-1, 10, 20, -5]})

def is_positive(grp_row):
    return grp_row["amount"].values > 0

df_result = df.groupby("id").apply(is_positive)

执行后df_result的输出结构为:

id   
1     False
      True
2     True
      False

由于函数限制无法使用Series操作,导致结果未保留原索引,需要将分组后的结果展开,使其索引/顺序与原df完全一致,确保df.iloc[i]与df_result.iloc[i]一一对应。


解决方案

方法1:使用groupby.transform(推荐)

transform会自动将分组计算结果按原DataFrame的索引对齐返回,只需保证函数返回与分组长度匹配的一维数组即可:

def is_positive(grp_row):
    return grp_row["amount"].values > 0

# 直接对amount列分组执行transform
df_result = df.groupby("id")["amount"].transform(is_positive)

输出结果:

0    False
1     True
2     True
3    False
Name: amount, dtype: bool

结果完全匹配原df的索引和顺序。

方法2:手动对齐apply结果

如果必须保留apply方式,有两种实现思路:

  1. 修改函数返回带原索引的Series:
def is_positive(grp_row):
    # 将计算结果与分组的原索引绑定
    return pd.Series(grp_row["amount"].values > 0, index=grp_row.index)

# 移除分组id的层级索引
df_result = df.groupby("id").apply(is_positive).reset_index(level=0, drop=True)
  1. 不修改函数,手动展开并映射原索引:
df_result = df.groupby("id").apply(is_positive)
# 展开嵌套结果,按原df的索引重新赋值
df_result = pd.Series(
    [val for sub_arr in df_result.values for val in sub_arr],
    index=df.index
)

两种方式最终都能得到与原df索引/顺序完全对齐的结果。


内容的提问来源于stack exchange,提问作者CutePoison

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 22:20:02