在函数中计算Pandas DataFrame加权平均值遇NameError问题求助
问题解决:DataFrame分组加权平均值计算的作用域错误
错误原因
你遇到的NameError: name 'df1' is not defined是因为weighted_mean函数试图访问some_function内部的局部变量df1,但两个函数不在同一作用域,weighted_mean无法获取到这个变量。
解决方案
下面提供两种可行的修正方式:
方式1:将加权平均函数嵌套在主函数内部
把weighted_mean定义在some_function里,这样它就能直接访问到df1变量:
import numpy as np import pandas as pd def some_function(df1=None): def weighted_mean(x): try: weights = df1.loc[x.index, 'amount'] return np.average(x, weights=weights) > 0.5 except ZeroDivisionError: return 0 df1 = df1.groupby('id').agg( xx=('amount', lambda x: x.sum() > 100), yy=('other_col', weighted_mean) ).reset_index() return df1 df2 = pd.DataFrame({'id':[1,1,2,2,3], 'amount':[10, 200, 1, 10, 150], 'other_col':[0.1, 0.6, 0.7, 0.2, 0.4]}) df2 = some_function(df1=df2) print(df2)
方式2:使用groupby.apply处理整个分组
这种方式更直观,直接对每个分组的全部列进行处理,不需要跨作用域访问变量:
import numpy as np import pandas as pd def process_group(group): xx = group['amount'].sum() > 100 try: weighted_avg = np.average(group['other_col'], weights=group['amount']) yy = weighted_avg > 0.5 except ZeroDivisionError: yy = 0 return pd.Series({'xx': xx, 'yy': yy}) def some_function(df1=None): df1 = df1.groupby('id').apply(process_group).reset_index() return df1 df2 = pd.DataFrame({'id':[1,1,2,2,3], 'amount':[10, 200, 1, 10, 150], 'other_col':[0.1, 0.6, 0.7, 0.2, 0.4]}) df2 = some_function(df1=df2) print(df2)
输出结果
两种方式都会得到你期望的输出:
id xx yy 0 1 True True 1 2 False False 2 3 True False
内容的提问来源于stack exchange,提问作者corianne1234
相关产品推荐
相关产品推荐

