You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas groupby后apply(lambda)报KeyError:找不到Industry_adjusted_return列

问题:GroupBy后调用Apply触发KeyError,找不到'Industry_adjusted_return'列

使用Python 3.8与pandas 1.3.3版本,对DataFrame执行groupby操作后调用.apply()计算Industry_adjusted_return列的加权平均时,触发KeyError,提示找不到该列。

复现代码

import pandas as pd

# 创建小型DataFrame
data = {
    'ISIN': ['DE000A1DAHH0', 'DE000KSAG888'],
    'Date': ['2017-03-01', '2017-03-01'],
    'MP_quintile': [0, 0],
    'Mcap_w': [8089460.00, 4154519.75],
    'Industry_adjusted_return': [-0.00869, 0.043052]
}

df = pd.DataFrame(data)
df['Date'] = pd.to_datetime(df['Date'])  # 确保Date为datetime类型

出错代码

# 注:原代码中wa未定义,应为df
for i, grouped in df.groupby(['Date','MP_quintile']):
     print(i, grouped)
     weighted_average_returns = grouped.apply(lambda x: (x['Industry_adjusted_return'] * (x['Mcap_w'] / x['Mcap_w'].sum())).sum())

报错信息

{
    "name": "KeyError",
    "message": "'Industry_adjusted_return'",
    "stack": "---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile c:\\Users\\mbkoo\\anaconda3\\envs\\myenv\\Lib\\site-packages\\pandas\\core\\indexes\\base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile c:\\Users\\mbkoo\\anaconda3\\envs\\myenv\\Lib\\site-packages\\pandas\\_libs\\index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile c:\\Users\\mbkoo\\anaconda3\\envs\\myenv\\Lib\\site-packages\\pandas\\_libs\\index.pyx:146, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas\\_libs\\index_class_helper.pxi:49, in pandas._libs.index.Int64Engine._check_type()\n\nKeyError: 'Industry_adjusted_return'\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[10], line 8\n      3  print(i,grouped)\n      4  #weighted_average_returns = grouped.apply( lambda x: ((x['Mcap_w'] / x['Mcap_w'].sum()))).sum()\n      5 # grouped['weights_EW'] = 1 / len(grouped)\n      6 # grouped['return_EW'] = grouped['Industry_adjusted_return'] * grouped['weights_EW']\n----> 8  weighted_average_returns = grouped.apply(lambda x: (x['Industry_adjusted_return'] * (x['Mcap_w'] / x['Mcap_w'].sum())).sum()) #\n      9 # equally_weighted_returns=grouped['return_EW'].sum()\n     10 # # _df=cpd.from_dataframe(_df,allow_copy=True)\n     11  break\n\nFile c:\\Users\\pandas\\core\\frame.py:9568, in DataFrame.apply(self, func, axis, raw, result_type, args, **kwargs)\n   9557 from pandas.core.apply import frame_apply\n   9559 op = frame_apply(\n   9560     self,\n   9561     func=func,\n   (...)\n   9566     kwargs=kwargs,\n   9567 )\n-> 9568 return op.apply().__finalize__(self, method=\"apply\")\n\nFile c:\\Users\\pandas\\core\\apply.py:764, in FrameApply.apply(self)\n    761 elif self.raw:\n    762     return self.apply_raw()\n-> 764 return self.apply_standard()\n\nFile c:\\Users\\pandas\\core\\apply.py:891, in FrameApply.apply_standard(self)\n    890 def apply_standard(self):\n-> 891     results, res_index = self.apply_series_generator()\n    893     # wrap results\n    894     return self.wrap_results(results, res_index)\n\nFile c:\\Users\\pandas\\core\\apply.py:907, in FrameApply.apply_series_generator(self)\n    904 with option_context(\"mode.chained_assignment\", None):\n    905     for i, v in enumerate(series_gen):\n    906         # ignore SettingWithCopy here in case the user mutates\n-> 907         results[i] = self.f(v)\n    908         if isinstance(results[i], ABCSeries):\n    909             # If we have a view on v, we need to make a copy because\n    910             #  series_generator will swap out the underlying data\n    911             results[i] = results[i].copy(deep=False)\n\nCell In[10], line 8, in <lambda>(x)\n      3  print(i,grouped)\n      4  #weighted_average_returns = grouped.apply( lambda x: ((x['Mcap_w'] / x['Mcap_w'].sum()))).sum()\n      5 # grouped['weights_EW'] = 1 / len(grouped)\n      6 # grouped['return_EW'] = grouped['Industry_adjusted_return'] * grouped['weights_EW']\n----> 8  weighted_average_returns = grouped.apply(lambda x: (x['Industry_adjusted_return'] * (x['Mcap_w'] / x['Mcap_w'].sum())).sum()) #\n      9 # equally_weighted_returns=grouped['return_EW'].sum()\n     10 # # _df=cpd.from_dataframe(_df,allow_copy=True)\n     11  break\n\nFile c:\\Users\\pandas\\core\\series.py:981, in Series.__getitem__(self, key)\n    978     return self._values[key]\n    980 elif key_is_scalar:\n-> 981     return self._get_value(key)\n    983 if is_hashable(key):\n    984     # Otherwise index.get_value will raise InvalidIndexError\n    985     try:\n    986         # For labels that don't resolve as scalars like tuples and frozensets\n\nFile c:\\Users\\pandas\\core\\series.py:1089, in Series._get_value(self, label, takeable)\n   1086     return self._values[label]\n   1088 # Similar to Index.get_value, but we do not fall back to positional\n-> 1089 loc = self.index.get_loc(label)\n   1090 return self.index._get_values_for_loc(self, loc, label)\n\nFile c:\\Users\\pandas\\core\\indexes\\base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 'Industry_adjusted_return'"
}

错误原因

  1. apply默认行为问题:对groupby得到的子DataFrame调用.apply()时,默认axis=0,会把每一列作为单独的Series传入lambda函数。此时lambda中的x是某一列的Series(比如ISIN列、Date列),你试图用x['Industry_adjusted_return']去取这个Series的索引值,自然找不到,触发KeyError。
  2. 代码笔误:原代码中wa.groupby未定义wa变量,应该是df.groupby。

修复方案

方案1:直接在groupby对象上调用apply(推荐)

不需要循环groupby结果,直接对groupby后的对象执行apply,此时传入lambda的是整个子DataFrame:

weighted_avg = df.groupby(['Date', 'MP_quintile']).apply(
    lambda x: (x['Industry_adjusted_return'] * (x['Mcap_w'] / x['Mcap_w'].sum())).sum()
)
print(weighted_avg)

方案2:用agg替代apply(性能更优)

避免apply的性能损耗,用agg结合自定义函数实现:

def weighted_average(group):
    weights = group['Mcap_w'] / group['Mcap_w'].sum()
    return (group['Industry_adjusted_return'] * weights).sum()

weighted_avg = df.groupby(['Date', 'MP_quintile']).agg(weighted_average)

方案3:用assign+eval简化计算

通过先计算权重列,再用eval完成加权求和:

weighted_avg = df.groupby(['Date', 'MP_quintile']).apply(
    lambda x: x.assign(weight=x['Mcap_w']/x['Mcap_w'].sum())
               .eval('Industry_adjusted_return * weight')
               .sum()
)

内容的提问来源于stack exchange,提问作者Mostafa Bouzari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 00:15:54