pandas groupby.apply弃用警告:添加include_groups=False后报错
问题说明
编写了一个Python函数,用于检查每个唯一DecId对应的Name唯一值数量:若一个DecId对应多个Name,则将Name追加到DecId后,确保每个Name与DecId的组合唯一。原代码可得到预期输出,但在Pandas 2.2版本中收到DeprecationWarning,尝试添加include_groups=False参数修复时触发KeyError: 'DecId'。
原代码
import pandas as pd # Sample DataFrame creation data = { 'DecId': ['D1', 'D1', 'D2', 'D2', 'D3', 'D3'], 'Name': ['JohnDoe1', 'JaneDoe2', 'AliceSmith', 'NON EEA', 'BobBrown', 'BobBrown'], 'OtherColumn': [10, 20, 30, 40, 50, 60] } df = pd.DataFrame(data) # Function to process each group and modify DecID if necessary due to multiple Names def mod_dec_ids_name(group): if len(group['Name'].unique()) > 1: #combine any instances of NON EEA into NON_EEA group['Name'] = group['Name'].str.replace('NON EEA', 'NON_EEA', regex=False) # If there are multiple unique Names, modify DecId group['DecId'] = group['DecId'] + '_' + group['Name'].apply(lambda x: x.split()[-1]) return group # Group by DecID and apply the processing function to check Names df2 = df.groupby('DecId').apply(lambda group: mod_dec_ids_name(group)).reset_index(drop=True) print(df2)
预期输出
DecId Name OtherColumn 0 D1_JohnDoe1 JohnDoe1 10 1 D1_JaneDoe2 JaneDoe2 20 2 D2_AliceSmith AliceSmith 30 3 D2_NON_EEA NON_EEA 40 4 D3 BobBrown 50 5 D3 BobBrown 60
收到的警告
:1: DeprecationWarning: DataFrameGroupBy.apply operated on the grouping columns. This behavior is deprecated, and in a future version of pandas the grouping columns will be excluded from the operation. Either pass
include_groups=Falseto exclude the groupings or explicitly select the grouping columns after groupby to silence this warning.
添加include_groups=False后的报错
df2 = df.groupby('DecId').apply(lambda group: mod_dec_ids_name(group), include_groups=False).reset_index(drop=True) Traceback (most recent call last): File "C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\indexes\base.py", line 3805, in get_loc return self._engine.get_loc(casted_key) File "index.pyx", line 167, in pandas._libs.index.IndexEngine.get_loc File "index.pyx", line 196, in pandas._libs.index.IndexEngine.get_loc File "pandas\_libs\hashtable_class_helper.pxi", line 7081, in pandas._libs.hashtable.PyObjectHashTable.get_item File "pandas\_libs\hashtable_class_helper.pxi", line 7089, in pandas._libs.hashtable.PyObjectHashTable.get_item KeyError: 'DecId' Traceback (most recent call last): File "<stdin>", line 1, in <module> File "C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\groupby\groupby.py", line 1819, in apply return self._python_apply_general(f, self._obj_with_exclusions) File "C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\groupby\groupby.py", line 1885, in _python_apply_general values, mutated = self._grouper.apply_groupwise(f, data, self.axis) res = f(group) File "<stdin>", line 1, in <lambda> File "<stdin>", line 4, in mod_dec_ids_name File "C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\frame.py", line 4102, in __getitem__ indexer = self.columns.get_loc(key) File "C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\indexes\base.py", line 3812, in get_loc raise KeyError(key) from err KeyError: 'DecId'
解决方案
方法1:使用include_groups=False并通过组名获取DecId
当设置include_groups=False时,分组列DecId不会包含在传入apply的子DataFrame中,此时可以通过group.name获取当前组的DecId值,修改处理函数如下:
import pandas as pd data = { 'DecId': ['D1', 'D1', 'D2', 'D2', 'D3', 'D3'], 'Name': ['JohnDoe1', 'JaneDoe2', 'AliceSmith', 'NON EEA', 'BobBrown', 'BobBrown'], 'OtherColumn': [10, 20, 30, 40, 50, 60] } df = pd.DataFrame(data) def mod_dec_ids_name(group): dec_id = group.name if len(group['Name'].unique()) > 1: group['Name'] = group['Name'].str.replace('NON EEA', 'NON_EEA', regex=False) # 使用group.name获取当前组的DecId group['DecId'] = dec_id + '_' + group['Name'].apply(lambda x: x.split()[-1]) else: # 单个Name时保持原DecId group['DecId'] = dec_id return group # 添加include_groups=False参数 df2 = df.groupby('DecId').apply(mod_dec_ids_name, include_groups=False).reset_index(drop=True) print(df2)
方法2:显式选择所有列消除警告
无需修改处理函数,只需在groupby后显式指定要保留的所有列,即可兼容旧行为并消除警告:
# 显式选择所有列,保持分组列在组内 df2 = df.groupby('DecId', as_index=False)[df.columns].apply(mod_dec_ids_name).reset_index(drop=True)
说明
- 方法1符合Pandas未来版本的行为,推荐使用;
- 方法2保留了旧版
groupby.apply的行为,无需修改原有函数,适合快速兼容。
内容的提问来源于stack exchange,提问作者botinky

