使用groupby apply时pd.Series.value_counts报错的问题
问题:Pandas GroupBy中调用apply+value_counts报错
问题场景
想用.apply(pd.Series.value_counts, axis=0)统计DataFrame中['a','b']两列的数值分布。在for循环遍历分组时代码能正常运行,但直接用groupby.apply却抛出错误:
TypeError: value_counts() got an unexpected keyword argument 'axis'
可复现代码
import pandas as pd import numpy as np # 示例数据框 df = pd.DataFrame( { "a": [1, 1, 2, 3, 3, 4, 5, 1, 1, 1, 4, 4, 4, 5, 6, 6, 6, 6, 3], 'b': [3, 4, 5, 5, 5, 2, 1, 3, 4, 4, 4, 5, 6, 6, 4, 3, 6, 6, 3], "Group": ['g1', 'g1', 'g1', 'g2', 'g2', 'g1', 'g2', 'g1', 'g1', 'g2','g2', 'g2', 'g2', 'g2', 'g1','g1', 'g2', 'g2', 'g2'], } ) # 正常运行的for循环分组代码 lst = [] for key, grp in df.groupby('Group'): df_ = grp[['a','b']].apply(pd.Series.value_counts, axis=0) df_['Group']=key lst.append(df_) print('正常输出结果:\n', pd.concat(lst)) # 报错的groupby.apply代码 df.groupby('Group')[['a','b']].apply(pd.Series.value_counts, axis=0)
预期输出
a b Group 1 4.0 NaN g1 2 1.0 1.0 g1 3 NaN 3.0 g1 4 1.0 3.0 g1 5 NaN 1.0 g1 6 2.0 NaN g1 1 1.0 1.0 g2 3 3.0 1.0 g2 4 3.0 2.0 g2 5 2.0 3.0 g2 6 2.0 4.0 g2
报错信息
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) /usr/local/lib/python3.7/dist-packages/pandas/core/groupby/groupby.py in apply(self, func, *args, **kwargs) 1274 try: -> 1275 result = self._python_apply_general(f, self._selected_obj) 1276 except TypeError: 11 frames TypeError: value_counts() got an unexpected keyword argument 'axis' During handling of the above exception, another exception occurred: TypeError Traceback (most recent call last) /usr/local/lib/python3.7/dist-packages/pandas/core/groupby/groupby.py in f(g) 1257 def f(g): 1258 with np.errstate(all="ignore"): -> 1259 return func(g, *args, **kwargs) 1260 1261 elif hasattr(nanops, "nan" + func): TypeError: value_counts() got an unexpected keyword argument 'axis'
原因分析
在for循环中,grp[['a','b']]是DataFrame,调用apply(pd.Series.value_counts, axis=0)时,axis=0是传给DataFrame.apply的参数,表示按列处理,每一列作为Series传入value_counts,和value_counts本身无关。
但在groupby.apply中,分组后的对象会被直接传给pd.Series.value_counts——而Series.value_counts没有axis参数,因此抛出错误。
解决方案
包装一层自定义函数,在函数内部对分组后的DataFrame执行apply(pd.Series.value_counts, axis=0),避免把axis参数直接传给value_counts:
自定义函数写法
def count_values(grp): return grp[['a','b']].apply(pd.Series.value_counts, axis=0) result = df.groupby('Group').apply(count_values).reset_index(level=0) result.rename(columns={'level_0':'Group'}, inplace=True) print(result)
Lambda简洁写法
result = df.groupby('Group')[['a','b']].apply(lambda x: x.apply(pd.Series.value_counts)).reset_index(level=0) result.rename(columns={'level_0':'Group'}, inplace=True) print(result)
验证结果
运行上述代码后,输出与for循环的预期结果完全一致。
内容的提问来源于stack exchange,提问作者Leo
相关产品推荐
相关产品推荐

