You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何以更简洁方式获取Pandas DataFrame重复元素的索引列表?

获取Pandas DataFrame指定列重复元素索引的更优方法

问题场景

需要提取Pandas DataFrame指定列中重复元素的索引列表,目前用的实现方式比较繁琐:

import pandas as pd
import numpy as np

test = ['a', 'a', 'b', 'c', 'b']
testdf = pd.DataFrame(test, columns=['test'])
np.asarray(np.where(list(testdf['test'].duplicated()))).tolist()[0]
# 输出结果: [1, 4]

更简洁的实现方法

方法1:直接利用布尔索引提取

duplicated()会返回一个布尔类型的Series,直接用它过滤DataFrame的索引,再转成列表即可,全程无需依赖numpy:

testdf[testdf['test'].duplicated()].index.tolist()
# 输出: [1, 4]

方法2:简化numpy写法

如果习惯用numpy,也可以省去多余的类型转换步骤,直接对duplicated()返回的Series调用np.where:

np.where(testdf['test'].duplicated())[0].tolist()
# 输出: [1, 4]

额外补充:获取所有重复元素的索引(含首次出现项)

如果你的需求是拿到所有重复元素的索引(比如'a'的0和1、'b'的2和4),可以给duplicated()传入keep=False参数:

testdf[testdf['test'].duplicated(keep=False)].index.tolist()
# 输出: [0, 1, 2, 4]

内容的提问来源于stack exchange,提问作者gaut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 19:40:57