如何在Pandas中按分组计算排除每组最后一行的累积乘积?
问题:按索引分组计算累积乘积(排除每组最后一个元素)
原始DataFrame
import pandas as pd n_obs = 3 dd = pd.DataFrame({ 'WTL_exploded': [0, 1, 2]*n_obs, 'hazard': [0.3, 0.4, 0.5, 0.2, 0.8, 0.9, 0.6,0.6,0.65], }, index=[1,1,1,2,2,2,3,3,3])
需求说明
按DataFrame的index分组,对每组的hazard列计算累积乘积,但仅保留每组中除最后一个元素之外的计算结果,最终期望输出如下:
| index | hazard |
|---|---|
| 1 | 0.3 |
| 1 | 0.12 |
| 2 | 0.2 |
| 2 | 0.16 |
| 3 | 0.6 |
| 3 | 0.36 |
解决方案
通过groupby分组后,对每组的hazard序列截取前n-1个元素(排除最后一个),再计算累积乘积,最后整理成目标格式:
# 分组处理并计算累积乘积 result = dd.groupby(level=0)['hazard'].apply(lambda x: x.iloc[:-1].cumprod()).reset_index() # 调整列名匹配期望输出 result.columns = ['index', 'hazard'] print(result)
代码解释
groupby(level=0):按DataFrame的最外层索引进行分组lambda x: x.iloc[:-1].cumprod():对每组的hazard列,用iloc[:-1]排除最后一个元素,再调用cumprod()计算累积乘积reset_index():将分组后的多级索引转换为普通列,得到与期望一致的结构
运行上述代码即可得到目标输出。
内容的提问来源于stack exchange,提问作者Luca Clissa
相关产品推荐
相关产品推荐

