使用group_by时遇TypeError:插入列与DataFrame索引不兼容
问题描述
从SQL读取的DataFrame结构如下:
id stock_id symbol date open high low close volume 0 1 35 ABSI 2022-09-28 3.06 3.33 3.0400 3.27 217040 1 2 35 ABSI 2022-09-29 3.19 3.19 3.0300 3.12 187309 2 3 35 ABSI 2022-09-30 3.11 3.27 3.0700 3.13 196566 3 4 35 ABSI 2022-10-03 3.16 3.16 2.8600 2.97 310441 4 5 35 ABSI 2022-10-04 3.04 3.37 2.9600 3.27 361082 .. ... ... ... ... ... ... ... ... ... 383 384 16 VVI 2022-10-03 31.93 33.85 31.3050 33.60 151357 384 385 16 VVI 2022-10-04 34.41 35.46 34.1900 35.39 105773 385 386 16 VVI 2022-10-05 34.67 35.30 34.5000 34.86 59605 386 387 16 VVI 2022-10-06 34.80 35.14 34.3850 34.50 55323 387 388 16 VVI 2022-10-07 33.99 33.99 33.3409 33.70 45187 [388 rows x 9 columns]
执行以下代码计算过去5天成交量均值并添加为新列时:
df['volume_5_day'] = df.groupby('stock_id')['volume'].rolling(5).mean()
抛出错误:
Traceback (most recent call last): File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/frame.py", line 11003, in _reindex_for_setitem reindexed_value = value.reindex(index)._values File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/series.py", line 4672, in reindex return super().reindex(**kwargs) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/generic.py", line 4966, in reindex return self._reindex_axes( File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/generic.py", line 4981, in _reindex_axes new_index, indexer = ax.reindex( File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/indexes/base.py", line 4237, in reindex target = self._wrap_reindex_result(target, indexer, preserve_names) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/indexes/multi.py", line 2520, in _wrap_reindex_result target = MultiIndex.from_tuples(target) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/indexes/multi.py", line 204, in new_meth return meth(self_or_cls, *args, **kwargs) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/indexes/multi.py", line 559, in from_tuples arrays = list(lib.tuples_to_object_array(tuples).T) File "pandas/_libs/lib.pyx", line 2930, in pandas._libs.lib.tuples_to_object_array ValueError: Buffer dtype mismatch, expected 'Python object' but got 'long' The above exception was the direct cause of the following exception: Traceback (most recent call last): File "/home/dan/Documents/code/wolfhound/add_indicators_daily.py", line 10, in <module> df['volume_10_day'] = df.groupby('stock_id')['volume'].rolling(1).mean() File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/frame.py", line 3655, in __setitem__ self._set_item(key, value) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/frame.py", line 3832, in _set_item value = self._sanitize_column(value) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/frame.py", line 4535, in _sanitize_column return _reindex_for_setitem(value, self.index) File "/home/dan/.local/lib/python3.10/site-packages/pandas/core/frame.py", line 11010, in _reindex_for_setitem raise TypeError( TypeError: incompatible index of inserted column with frame index
该代码之前可正常运行,现在报错,询问问题原因及解决方法。
问题原因及解决方法
错误原因
groupby后调用rolling返回的结果带有多层索引(MultiIndex),外层是stock_id分组键,内层是原DataFrame的索引。而原DataFrame使用单层索引,直接赋值会导致索引不匹配,触发报错。
之前能运行大概率是pandas版本差异导致——旧版本可能自动处理了索引对齐逻辑,新版本则严格校验了索引结构。
解决方法
需要将滚动计算的结果重置索引,去掉分组带来的额外索引层,让结果索引和原DataFrame对齐,两种常用方案:
方案1:用reset_index(drop=True)处理索引
df['volume_5_day'] = df.groupby('stock_id')['volume'].rolling(5).mean().reset_index(drop=True)
drop=True会直接丢弃分组产生的额外索引层,让结果回到单层索引,和原DataFrame匹配。
方案2:用transform替代直接赋值
transform会自动保持原DataFrame的索引结构,无需手动处理索引:
df['volume_5_day'] = df.groupby('stock_id')['volume'].transform(lambda x: x.rolling(5).mean())
额外注意事项
如果数据未按stock_id和date排序,建议先执行排序,避免滚动计算逻辑错误:
df = df.sort_values(['stock_id', 'date'])
内容的提问来源于stack exchange,提问作者a7dc
相关产品推荐
相关产品推荐

