You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中筛选DataFrame某列各值的前x次出现?

高效筛选Pandas DataFrame中每个值的前x次出现

直接用Pandas的分组计数功能就能搞定,完全不用低效的iterrows(),核心是矢量化操作,处理大数据集效率拉满:

  1. 给每个分组添加出现序号
    用groupby()按目标列分组,再通过cumcount()为每个组内的行生成从0开始的连续序号,这个序号就是当前值在组内的出现顺序。

  2. 筛选前x次出现的行
    只保留组内序号小于x的行,就能得到每个值的前x次出现结果。

代码示例

假设你的DataFrame是df,目标列名为target_col,要保留每个值的前3次出现:

import pandas as pd

# 构造示例数据
data = {
    'target_col': ['A', 'A', 'B', 'A', 'B', 'B', 'C', 'C', 'A', 'C'],
    'value': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
}
df = pd.DataFrame(data)

x = 3  # 指定要保留的前x次出现
# 执行筛选
filtered_df = df[df.groupby('target_col').cumcount() < x]

print(filtered_df)

输出结果

target_col  value
0          A      1
1          A      2
2          B      3
3          A      4
4          B      5
5          B      6
6          C      7
7          C      8
9          C     10

可以看到每个值都精准保留了前3次出现的行,全程没有循环,效率比iterrows()高几个量级。

内容的提问来源于stack exchange,提问作者Raido Lacorte Galina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 18:57:03