如何在Python中筛选提交数据集中被Stable包裹的含Increase/Decrease的行范围
问题描述
现有如下提交记录数据集:
| commit | bug |
|---|---|
| sha_1 | Stable |
| sha_2 | Stable |
| sha_3 | Stable |
| sha_4 | Increase |
| sha_5 | Stable |
| sha_6 | Stable |
| sha_7 | Decrease |
| sha_8 | Stable |
| sha_9 | Decrease |
| sha_10 | Decrease |
| sha_11 | Increase |
| sha_12 | Stable |
需要筛选出同时包含Increase和Decrease(任意顺序)、且被两条Stable提交包裹的连续行范围。根据示例,预期输出如下(注:示例中重复的sha_8为笔误,实际应为连续的sha_8到sha_12):
| commit | bug |
|---|---|
| sha_3 | Stable |
| sha_4 | Increase |
| sha_5 | Stable |
| sha_6 | Stable |
| sha_7 | Decrease |
| sha_8 | Stable |
| sha_9 | Decrease |
| sha_10 | Decrease |
| sha_11 | Increase |
| sha_12 | Stable |
可行解决方案
方案一:使用Pandas处理(适合大型数据集)
利用Pandas的分组和索引功能,快速定位符合条件的行区间:
import pandas as pd # 构造数据集(实际场景可从文件读取) data = { 'commit': ['sha_1', 'sha_2', 'sha_3', 'sha_4', 'sha_5', 'sha_6', 'sha_7', 'sha_8', 'sha_9', 'sha_10', 'sha_11', 'sha_12'], 'bug': ['Stable', 'Stable', 'Stable', 'Increase', 'Stable', 'Stable', 'Decrease', 'Stable', 'Decrease', 'Decrease', 'Increase', 'Stable'] } df = pd.DataFrame(data) # 标记非Stable行,并生成连续非稳定块的ID df['is_non_stable'] = df['bug'].isin(['Increase', 'Decrease']) df['block_id'] = (df['is_non_stable'] != df['is_non_stable'].shift()).cumsum() # 分组统计每个非稳定块的关键信息 blocks = df[df['is_non_stable']].groupby('block_id').agg( has_increase=('bug', lambda x: 'Increase' in x.values), has_decrease=('bug', lambda x: 'Decrease' in x.values), start_idx=('bug', 'idxmin'), end_idx=('bug', 'idxmax') ).reset_index() # 筛选同时包含Increase和Decrease的有效块 valid_blocks = blocks[(blocks['has_increase'] & blocks['has_decrease'])] # 收集需要保留的行索引 keep_indices = set() for _, row in valid_blocks.iterrows(): start_idx = row['start_idx'] end_idx = row['end_idx'] # 找到块前最近的Stable行 prev_stable_idx = df.loc[:start_idx-1][df['bug'] == 'Stable'].index.max() # 找到块后最近的Stable行 next_stable_idx = df.loc[end_idx+1:][df['bug'] == 'Stable'].index.min() # 添加整个区间的索引 keep_indices.update(range(prev_stable_idx, next_stable_idx + 1)) # 生成结果并重置索引 result_df = df.loc[sorted(keep_indices)].reset_index(drop=True) # 输出结果(仅保留commit和bug列) print(result_df[['commit', 'bug']])
逻辑说明
- 标记所有非Stable行,通过连续变化生成块ID,将连续的非Stable行归为同一组。
- 对每个块检查是否同时包含Increase和Decrease,筛选出有效块。
- 为每个有效块扩展范围:向前取最近的Stable行,向后取最近的Stable行,保留整个区间的所有行。
方案二:纯Python实现(无第三方依赖,适合小型数据集)
直接遍历列表,定位符合条件的区间:
# 原始数据集 data = [ ('sha_1', 'Stable'), ('sha_2', 'Stable'), ('sha_3', 'Stable'), ('sha_4', 'Increase'), ('sha_5', 'Stable'), ('sha_6', 'Stable'), ('sha_7', 'Decrease'), ('sha_8', 'Stable'), ('sha_9', 'Decrease'), ('sha_10', 'Decrease'), ('sha_11', 'Increase'), ('sha_12', 'Stable'), ] valid_ranges = [] current_start = None current_types = set() # 遍历数据集,定位有效非稳定区间 for idx, (commit, bug) in enumerate(data): if bug in ['Increase', 'Decrease']: if current_start is None: current_start = idx current_types.add(bug) else: if current_start is not None: # 检查当前区间是否包含两种类型 if len(current_types) == 2: # 向前查找最近的Stable行 prev_stable_idx = idx - 1 while prev_stable_idx >= 0 and data[prev_stable_idx][1] != 'Stable': prev_stable_idx -= 1 # 向后的Stable行就是当前idx valid_ranges.append((prev_stable_idx, idx)) # 重置当前区间状态 current_start = None current_types = set() # 收集结果 result = [] for start, end in valid_ranges: result.extend(data[start:end+1]) # 打印结果 for item in result: print(f"{item[0]} | {item[1]}")
逻辑说明
- 遍历过程中记录连续的非Stable区间,同时跟踪区间内的类型。
- 当遇到Stable行时,检查之前的非Stable区间是否包含两种类型,若是则扩展区间到前后的Stable行。
- 最后收集所有有效区间的行,输出结果。
内容的提问来源于stack exchange,提问作者Giammaria GIORDANO
相关产品推荐
相关产品推荐

