按Symbol分组添加EMA触发ValueError:无法在重复标签轴重索引
问题:分组计算EMA指标时触发「cannot reindex on an axis with duplicate labels」报错
数据示例
timestamp open high low close volume trade_count vwap symbol volume_10_day 0 2021-06-28 11:15:00+00:00 1.8899 1.89 1.8899 1.8900 1200 2 1.889942 AADI 102027.6 1 2021-06-28 13:15:00+00:00 1.8200 1.82 1.8100 1.8100 710 8 1.818282 AADI 102027.6 2 2021-06-28 13:30:00+00:00 1.8600 1.87 1.8500 1.8600 13845 52 1.858409 AADI 102027.6 3 2021-06-28 13:35:00+00:00 1.8659 1.87 1.8626 1.8700 860 5 1.867670 AADI 102027.6 4 2021-06-28 13:40:00+00:00 1.8700 1.87 1.8477 1.8477 7386 33 1.859019 AADI 102027.6 ... ... ... ... ... ... ... ... ... ... ... 423069 2021-01-13 00:35:00+00:00 1.2600 1.27 1.2500 1.2600 242391 140 1.260218 ZOM 122065.8 423070 2021-01-13 00:40:00+00:00 1.2600 1.26 1.2500 1.2600 129074 108 1.256255 ZOM 122065.8 423071 2021-01-13 00:45:00+00:00 1.2500 1.26 1.2500 1.2500 297198 151 1.253695 ZOM 122065.8 423072 2021-01-13 00:50:00+00:00 1.2600 1.26 1.2500 1.2600 223822 121 1.256325 ZOM 122065.8 423073 2021-01-13 00:55:00+00:00 1.2600 1.26 1.2500 1.2600 378248 222 1.255110 ZOM 122065.8 [423074 rows x 10 columns]
尝试的代码
ema72 = lambda x: ta.ema(df.loc[x.index, "close"], 72) ema89 = lambda x: ta.ema(df.loc[x.index, "close"], 89) ema216 = lambda x: ta.ema(df.loc[x.index, "close"], 216) ema267 = lambda x: ta.ema(df.loc[x.index, "close"], 267) ema200 = lambda x: ta.ema(df.loc[x.index, "close"], 200) df["EMA72"] = df.groupby(['symbol']).apply(ema72).reset_index(0,drop=True) df["EMA89"] = df.groupby(['symbol']).apply(ema89).reset_index(0,drop=True) df["EMA216"] = df.groupby(['symbol']).apply(ema216).reset_index(0,drop=True) df["EMA267"] = df.groupby(['symbol']).apply(ema267).reset_index(0,drop=True) df["EMA200"] = df.groupby(['symbol']).apply(ema200).reset_index(0,drop=True)
报错信息
raise ValueError("cannot reindex on an axis with duplicate labels") ValueError: cannot reindex on an axis with duplicate labels
问题原因及解决办法
1. 核心错误点
你的lambda函数错误地引用了整个DataFrame的索引(df.loc[x.index, "close"]),而非分组后的子数据集x。这会导致apply返回的结果索引与原DataFrame的索引重叠冲突,最终触发reindex失败。
另外,groupby.apply返回的结果默认带有多层索引,即使执行reset_index(0, drop=True),若原DataFrame本身存在重复索引,依然会导致对齐失败。
2. 正确解决方案
步骤1:确保原数据索引唯一
先重置索引,避免原索引存在重复值:
df = df.reset_index(drop=True)
步骤2:用groupby.transform替代apply
transform会自动返回与原数据长度一致的结果,并自动对齐索引,无需手动处理索引问题。直接基于分组后的子数据计算EMA:
import talib as ta # 分组计算各周期EMA df["EMA72"] = df.groupby('symbol')['close'].transform(lambda x: ta.ema(x, 72)) df["EMA89"] = df.groupby('symbol')['close'].transform(lambda x: ta.ema(x, 89)) df["EMA216"] = df.groupby('symbol')['close'].transform(lambda x: ta.ema(x, 216)) df["EMA267"] = df.groupby('symbol')['close'].transform(lambda x: ta.ema(x, 267)) df["EMA200"] = df.groupby('symbol')['close'].transform(lambda x: ta.ema(x, 200))
3. 滚动均值的正确写法
若要计算滚动均值,同样用transform:
df["MA200"] = df.groupby('symbol')['close'].transform(lambda x: x.rolling(200).mean())
或者用rolling后重置索引(需确保原索引唯一):
df["MA200"] = df.groupby('symbol')['close'].rolling(200).mean().reset_index(0, drop=True)
内容的提问来源于stack exchange,提问作者a7dc
相关产品推荐
相关产品推荐

