You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效删除DataFrame中不符合1-2-3-4循环序列的行?

高效筛选DataFrame中符合1-2-3-4循环序列的行

我需要处理一个大型DataFrame,要求只保留B列值严格符合1-2-3-4-1-2-3……循环序列的行,删除所有不符合的行。

现有数据示例

A     B
12/2/2022    0.02   2
14/2/2022    0.01   1
15/2/2022    0.04   4
16/2/2022   -0.02   3
18/2/2022   -0.01   2
20/2/2022    0.04   1
21/2/2022    0.02   3
22/2/2022   -0.01   1
24/2/2022    0.04   4
26/2/2022   -0.02   2
27/2/2022    0.01   3
28/2/2022    0.04   1
01/3/2022   -0.02   3
03/3/2022   -0.01   2
05/3/2022    0.04   1
06/3/2022    0.02   3
08/3/2022   -0.01   1
10/3/2022    0.04   4
12/3/2022   -0.02   2
13/3/2022    0.01   3
15/3/2022    0.04   1
      ...

期望结果示例

A     B
14/2/2022    0.01   1
18/2/2022   -0.01   2
21/2/2022    0.02   3
24/2/2022    0.04   4
28/2/2022    0.04   1
03/3/2022   -0.01   2
06/3/2022    0.02   3
10/3/2022    0.04   4
15/3/2022    0.04   1
        ...

当前方案的问题

我现在用4个类似以下的循环逐个检查序列衔接(4→1、1→2、2→3、3→4),这种方法繁琐且低效,处理大数据时速度很慢:

df_len = len(df)
df_len2 = 0
while df_len != df_len2:
    df_len = len(df)
    df.loc[(df.B.shift(1) == 4) & (df.B != 1), 'B'] = 0
    df = df[df['B'] != 0]
    df_len2 = len(df)

基于NumPy的高效解决方案

核心思路是预先记录每个目标值(1、2、3、4)的所有出现位置,然后用二分查找快速定位下一个符合序列要求的值的位置,全程用NumPy操作避免冗余循环,大幅提升效率。

import numpy as np
import pandas as pd

# 假设你的DataFrame是df
b_vals = df['B'].to_numpy()
indices = df.index.to_numpy()

# 记录每个值对应的索引列表
value_positions = {
    1: indices[b_vals == 1],
    2: indices[b_vals == 2],
    3: indices[b_vals == 3],
    4: indices[b_vals == 4]
}

# 初始化要保留的索引列表,从第一个1开始
keep_indices = []
current_target = 1
last_idx = -np.inf

while True:
    # 用二分查找找到第一个大于last_idx的当前目标值的索引位置
    pos_list = value_positions[current_target]
    idx = np.searchsorted(pos_list, last_idx, side='right')
    if idx >= len(pos_list):
        break  # 没有更多符合条件的值,退出循环
    
    selected_idx = pos_list[idx]
    keep_indices.append(selected_idx)
    # 更新下一个目标值:1→2→3→4→1...
    current_target = current_target % 4 + 1
    last_idx = selected_idx

# 筛选DataFrame
filtered_df = df.loc[keep_indices]

方案说明

  1. 预存位置:先把每个值(1-4)在DataFrame中的索引都收集起来,这一步是O(n)的线性操作,非常高效。
  2. 二分查找:用np.searchsorted快速定位下一个符合要求的索引,时间复杂度是O(log k)(k是当前值的出现次数),比逐行循环快几个数量级。
  3. 无循环冗余:整个逻辑只需要一轮循环,循环次数等于最终保留的行数,远小于原数据行数,适合处理大型数据集。

内容的提问来源于stack exchange,提问作者Gus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:50:21