如何在Pandas中处理含零值的百分比变化计算?
解决分组计算百分比变化时除以零的问题
你的需求是对分组后的score列计算百分比变化,同时处理前值为0的场景——当从0变为非0值时,让百分比变化返回100%(即1.0),而非默认的inf。
方法一:自定义百分比变化函数
可以编写自定义函数覆盖默认的pct_change()逻辑,针对性处理除以零的情况:
import pandas as pd df = pd.DataFrame({'team': ['A', 'A', 'A', 'B', 'B', 'B', 'C', 'C'], 'points': [12, 0, 19, 22, 0, 25, 0, 30], 'score': [12, 0, 19, 22, 0, 25, 0, 30] }) def custom_pct_change(series): diff = series.diff() # 计算当前值与前值的差值 prev_values = series.shift(1) # 获取前一个值 # 遍历每一行:前值不为0时用常规公式,前值为0时返回1.0 return pd.Series( [(d / p) if p != 0 else 1.0 for d, p in zip(diff, prev_values)], index=series.index ) # 分组应用自定义函数 df['score'] = df.groupby('team', sort=False)['score'].apply(custom_pct_change).to_numpy() print(df)
方法二:用where简化逻辑
如果想要更简洁的代码,可以基于默认的pct_change()结果,用where方法替换前值为0时的inf:
import pandas as pd df = pd.DataFrame({'team': ['A', 'A', 'A', 'B', 'B', 'B', 'C', 'C'], 'points': [12, 0, 19, 22, 0, 25, 0, 30], 'score': [12, 0, 19, 22, 0, 25, 0, 30] }) df['score'] = df.groupby('team', sort=False)['score'].apply( lambda x: x.pct_change().where(x.shift(1) != 0, 1.0) ).to_numpy() print(df)
运行结果
两种方法都会得到相同的输出:
team points score 0 A 12 NaN 1 A 0 -1.0 2 A 19 1.0 3 B 22 NaN 4 B 0 -1.0 5 B 25 1.0 6 C 0 NaN 7 C 30 1.0
- 每组第一个值因无前置值保留
NaN(与pct_change()默认行为一致) - 从非0值变为0时,按常规逻辑计算为-100%(即-1.0)
- 从0变为非0值时,返回1.0(100%),完全匹配你的需求
内容的提问来源于stack exchange,提问作者Bad Coder
相关产品推荐
相关产品推荐

