Pandas groupby.apply弃用警告消除方法咨询
问题背景
运行读取CSV数据的Python脚本时,触发了如下弃用警告:
DeprecationWarning: DataFrameGroupBy.apply operated on the grouping columns. This behavior is deprecated, and in a future version of pandas the grouping columns will be excluded from the operation. Either pass
include_groups=Falseto exclude the groupings or explicitly select the grouping columns after groupby to silence this warning.
警告源于这段代码:
fprice = df.groupby(['StartDate', 'Commodity', 'DealType']).apply(lambda group: -(group['MTMValue'].sum() - (group['FixedPriceStrike'] * group['Quantity']).sum()) / group['Quantity'].sum()).reset_index(name='FloatPrice')
原因分析
这个警告是因为pandas即将调整GroupBy.apply的默认行为:当前版本中,分组后的DataFrame会包含分组列,但未来版本会默认排除分组列。警告要求你明确指定是否要在apply操作中包含分组列。你的lambda函数并未使用分组列,只需明确告知pandas排除分组列即可消除警告。
解决方案
有两种简单的修复方式:
方案一:添加include_groups=False参数
在apply方法中传入该参数,明确告知pandas不要将分组列传入lambda函数,这与未来版本的默认行为一致:
fprice = df.groupby(['StartDate', 'Commodity', 'DealType']).apply( lambda group: -(group['MTMValue'].sum() - (group['FixedPriceStrike'] * group['Quantity']).sum()) / group['Quantity'].sum(), include_groups=False ).reset_index(name='FloatPrice')
方案二:显式选择需要的列
在groupby之后直接指定lambda函数用到的列,这样分组后的DataFrame只包含这些列,自然不会触发警告:
fprice = df.groupby(['StartDate', 'Commodity', 'DealType'])[['MTMValue', 'FixedPriceStrike', 'Quantity']].apply( lambda group: -(group['MTMValue'].sum() - (group['FixedPriceStrike'] * group['Quantity']).sum()) / group['Quantity'].sum() ).reset_index(name='FloatPrice')
验证
两种方案都能生成你预期的输出(为每行添加FloatPrice列),且不会再触发弃用警告。
内容的提问来源于stack exchange,提问作者iBeMeltin

