You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas Groupby按组基于指定列值生成新列

问题描述

现有如下DataFrame:

import pandas as pd

df = pd.DataFrame({'c0':['1980']*3+['1990']*2+['2000']*3,
                   'c1':['x','y','z']+['x','y']+['x','y','z'],
                   'c2':range(8)  })

初始表格输出:

c0 c1  c2
0  1980  x   0
1  1980  y   1
2  1980  z   2
3  1990  x   3
4  1990  y   4
5  2000  x   5
6  2000  y   6
7  2000  z   7

需求:通过Pandas的groupby按c0(年份列)分组,每组内将c2的所有值减去该组中c1为'y'对应的c2值,生成新列c3,预期结果如下:

c0 c1  c2  c3
0  1980  x   0  -1
1  1980  y   1   0
2  1980  z   2   1
3  1990  x   3  -1
4  1990  y   4   0
5  2000  x   5  -1
6  2000  y   6   0
7  2000  z   7   1

目前已通过循环实现功能,但代码不够简洁,尝试groupby的transform和apply方法未成功,寻求优化方案。

优化方案

方法1:利用groupby.transform直接计算

通过分组后提取每组中c1=='y'的c2值,再用transform保证结果长度与原DataFrame匹配,最后做差值计算:

df['c3'] = df['c2'] - df.groupby('c0').transform(lambda grp: grp[grp.index.isin(df[df['c1'] == 'y'].index)].iloc[0])

方法2:构建基准值映射字典(性能更优)

先提取每个年份对应的'y'的c2值,生成映射字典,再通过map匹配后批量计算,适合大数据量场景:

# 生成年份到对应y值的映射
y_base = df[df['c1'] == 'y'].set_index('c0')['c2'].to_dict()
# 计算c3列
df['c3'] = df['c2'] - df['c0'].map(y_base)

方法3:结合groupby.apply与reindex

通过分组提取每组的基准值,再重新索引匹配原DataFrame的年份列,最后做差值:

base_values = df.groupby('c0').apply(lambda x: x.loc[x['c1'] == 'y', 'c2'].iloc[0])
df['c3'] = df['c2'] - base_values.reindex(df['c0']).values

以上三种方法均可得到预期结果,其中方法2的执行效率最高。

内容的提问来源于stack exchange,提问作者Spring

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:55:03