如何高效删除DataFrame中不符合1-2-3-4循环序列的行?
高效筛选DataFrame中符合1-2-3-4循环序列的行
我需要处理一个大型DataFrame,要求只保留B列值严格符合1-2-3-4-1-2-3……循环序列的行,删除所有不符合的行。
现有数据示例
A B 12/2/2022 0.02 2 14/2/2022 0.01 1 15/2/2022 0.04 4 16/2/2022 -0.02 3 18/2/2022 -0.01 2 20/2/2022 0.04 1 21/2/2022 0.02 3 22/2/2022 -0.01 1 24/2/2022 0.04 4 26/2/2022 -0.02 2 27/2/2022 0.01 3 28/2/2022 0.04 1 01/3/2022 -0.02 3 03/3/2022 -0.01 2 05/3/2022 0.04 1 06/3/2022 0.02 3 08/3/2022 -0.01 1 10/3/2022 0.04 4 12/3/2022 -0.02 2 13/3/2022 0.01 3 15/3/2022 0.04 1 ...
期望结果示例
A B 14/2/2022 0.01 1 18/2/2022 -0.01 2 21/2/2022 0.02 3 24/2/2022 0.04 4 28/2/2022 0.04 1 03/3/2022 -0.01 2 06/3/2022 0.02 3 10/3/2022 0.04 4 15/3/2022 0.04 1 ...
当前方案的问题
我现在用4个类似以下的循环逐个检查序列衔接(4→1、1→2、2→3、3→4),这种方法繁琐且低效,处理大数据时速度很慢:
df_len = len(df) df_len2 = 0 while df_len != df_len2: df_len = len(df) df.loc[(df.B.shift(1) == 4) & (df.B != 1), 'B'] = 0 df = df[df['B'] != 0] df_len2 = len(df)
基于NumPy的高效解决方案
核心思路是预先记录每个目标值(1、2、3、4)的所有出现位置,然后用二分查找快速定位下一个符合序列要求的值的位置,全程用NumPy操作避免冗余循环,大幅提升效率。
import numpy as np import pandas as pd # 假设你的DataFrame是df b_vals = df['B'].to_numpy() indices = df.index.to_numpy() # 记录每个值对应的索引列表 value_positions = { 1: indices[b_vals == 1], 2: indices[b_vals == 2], 3: indices[b_vals == 3], 4: indices[b_vals == 4] } # 初始化要保留的索引列表,从第一个1开始 keep_indices = [] current_target = 1 last_idx = -np.inf while True: # 用二分查找找到第一个大于last_idx的当前目标值的索引位置 pos_list = value_positions[current_target] idx = np.searchsorted(pos_list, last_idx, side='right') if idx >= len(pos_list): break # 没有更多符合条件的值,退出循环 selected_idx = pos_list[idx] keep_indices.append(selected_idx) # 更新下一个目标值:1→2→3→4→1... current_target = current_target % 4 + 1 last_idx = selected_idx # 筛选DataFrame filtered_df = df.loc[keep_indices]
方案说明
- 预存位置:先把每个值(1-4)在DataFrame中的索引都收集起来,这一步是O(n)的线性操作,非常高效。
- 二分查找:用
np.searchsorted快速定位下一个符合要求的索引,时间复杂度是O(log k)(k是当前值的出现次数),比逐行循环快几个数量级。 - 无循环冗余:整个逻辑只需要一轮循环,循环次数等于最终保留的行数,远小于原数据行数,适合处理大型数据集。
内容的提问来源于stack exchange,提问作者Gus
相关产品推荐
相关产品推荐

