You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas筛选指定列中连续重复指定值的行

筛选Pandas DataFrame中列x连续重复指定值(3和5)的行

需求说明

需要从大型Pandas DataFrame中,筛选出列x里连续重复的3或5对应的行,同时添加consecutive-count列标记同一连续组的编号。

示例输入

col row x   y
1   1   1   1
5   7   3   0
2   2   2   2
6   3   3   8
9   2   3   4
5   3   3   9
4   9   4   4
5   5   5   1
3   7   5   2
6   6   6   6
5   8   6   2
3   7   6   0

期望输出

col row x   y   consecutive-count
6   3   3   8          1
9   2   3   4          1 
5   3   3   9          1
5   5   5   1          2
3   7   5   2          2 

已尝试方法的问题

  • 方法1:
    m = df['x'].eq(df['x'].shift())
    df[m|m.shift(-1, fill_value=False)]
    
    问题:会把连续重复的6也包含进来,不符合仅保留3和5的要求。
  • 方法2:
    df.query('x in [3,5]')
    
    问题:会输出所有x为3或5的行,包括单独出现的(比如输入里第一行的3),而不是仅连续重复的行。

解决方案

通过分组标记连续重复组,再筛选符合条件的组即可实现需求:

  1. 生成连续相同值的分组键:
    df['group'] = (df['x'] != df['x'].shift()).cumsum()
    
  2. 统计每个分组的行数和对应x值,筛选出有效分组:
    group_stats = df.groupby('group').agg(
        size=('x', 'size'),
        x_val=('x', 'first')
    )
    valid_groups = group_stats[(group_stats['size'] >= 2) & (group_stats['x_val'].isin([3,5]))].index
    
  3. 过滤有效分组的行,生成最终的consecutive-count列:
    result = df[df['group'].isin(valid_groups)].copy()
    result['consecutive-count'] = result['group'].rank(method='dense').astype(int)
    result = result.drop('group', axis=1)
    

运行上述代码后,就能得到符合期望的输出结果。

内容的提问来源于stack exchange,提问作者code_error

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 05:06:43