如何用Pandas高效删除首个连续stage为0的分组行?
高效删除DataFrame开头连续指定值的行
需求说明
需要删除DataFrame开头连续满足stage=0的所有行,示例如下:
原表格:
| stage | h1 | h2 | h3 |
|---|---|---|---|
| 0 | 4 | 55 | 55 |
| 0 | 5 | 66 | 44 |
| 0 | 4 | 66 | 33 |
| 1 | 3 | 33 | 55 |
| 0 | 5 | 44 | 33 |
处理后表格:
| stage | h1 | h2 | h3 |
|---|---|---|---|
| 1 | 3 | 33 | 55 |
| 0 | 5 | 44 | 33 |
原实现用Python循环遍历,大数据量下性能不足,以下是几种原生Pandas高效解法:
解法1:利用cummax过滤
cummax()会将第一个非0值之后的所有行标记为True,直接过滤即可:
import pandas as pd data = {'stage': [0, 0, 0, 1, 0], 'h1': [4, 5, 4, 3, 5], 'h2': [55, 66, 66, 33, 44], 'h3': [55, 44, 33, 55, 33]} df = pd.DataFrame(data) # 过滤掉开头连续的0 df_filtered = df[(df['stage'] != 0).cummax()] df_filtered = df_filtered.reset_index(drop=True) print(df_filtered)
解法2:定位第一个非0行切片
找到第一个stage≠0的索引,直接从该位置开始切片:
import pandas as pd data = {'stage': [0, 0, 0, 1, 0], 'h1': [4, 5, 4, 3, 5], 'h2': [55, 66, 66, 33, 44], 'h3': [55, 44, 33, 55, 33]} df = pd.DataFrame(data) # 获取第一个非0行的索引 first_non_zero_idx = (df['stage'] != 0).idxmax() # 切片并重置索引 df_filtered = df.loc[first_non_zero_idx:].reset_index(drop=True) print(df_filtered)
解法3:利用cumsum过滤
通过cumsum()将第一个非0之后的行标记为非0值,再筛选:
import pandas as pd data = {'stage': [0, 0, 0, 1, 0], 'h1': [4, 5, 4, 3, 5], 'h2': [55, 66, 66, 33, 44], 'h3': [55, 44, 33, 55, 33]} df = pd.DataFrame(data) # 生成筛选条件 filter_condition = (df['stage'] != 0).cumsum() != 0 df_filtered = df[filter_condition].reset_index(drop=True) print(df_filtered)
性能优势
以上方法均为Pandas原生向量运算,避免了Python层面的循环遍历,在大数据量场景下(比如百万级行),速度会比原循环实现快数十倍甚至上百倍。
内容的提问来源于stack exchange,提问作者Akshay
相关产品推荐
相关产品推荐

