You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中筛选提交数据集中被Stable包裹的含Increase/Decrease的行范围

问题描述

现有如下提交记录数据集:

commitbug
sha_1Stable
sha_2Stable
sha_3Stable
sha_4Increase
sha_5Stable
sha_6Stable
sha_7Decrease
sha_8Stable
sha_9Decrease
sha_10Decrease
sha_11Increase
sha_12Stable

需要筛选出同时包含Increase和Decrease(任意顺序)、且被两条Stable提交包裹的连续行范围。根据示例,预期输出如下(注:示例中重复的sha_8为笔误,实际应为连续的sha_8到sha_12):

commitbug
sha_3Stable
sha_4Increase
sha_5Stable
sha_6Stable
sha_7Decrease
sha_8Stable
sha_9Decrease
sha_10Decrease
sha_11Increase
sha_12Stable
可行解决方案

方案一:使用Pandas处理(适合大型数据集)

利用Pandas的分组和索引功能,快速定位符合条件的行区间:

import pandas as pd

# 构造数据集(实际场景可从文件读取)
data = {
    'commit': ['sha_1', 'sha_2', 'sha_3', 'sha_4', 'sha_5', 'sha_6', 'sha_7', 'sha_8', 'sha_9', 'sha_10', 'sha_11', 'sha_12'],
    'bug': ['Stable', 'Stable', 'Stable', 'Increase', 'Stable', 'Stable', 'Decrease', 'Stable', 'Decrease', 'Decrease', 'Increase', 'Stable']
}
df = pd.DataFrame(data)

# 标记非Stable行,并生成连续非稳定块的ID
df['is_non_stable'] = df['bug'].isin(['Increase', 'Decrease'])
df['block_id'] = (df['is_non_stable'] != df['is_non_stable'].shift()).cumsum()

# 分组统计每个非稳定块的关键信息
blocks = df[df['is_non_stable']].groupby('block_id').agg(
    has_increase=('bug', lambda x: 'Increase' in x.values),
    has_decrease=('bug', lambda x: 'Decrease' in x.values),
    start_idx=('bug', 'idxmin'),
    end_idx=('bug', 'idxmax')
).reset_index()

# 筛选同时包含Increase和Decrease的有效块
valid_blocks = blocks[(blocks['has_increase'] & blocks['has_decrease'])]

# 收集需要保留的行索引
keep_indices = set()
for _, row in valid_blocks.iterrows():
    start_idx = row['start_idx']
    end_idx = row['end_idx']
    # 找到块前最近的Stable行
    prev_stable_idx = df.loc[:start_idx-1][df['bug'] == 'Stable'].index.max()
    # 找到块后最近的Stable行
    next_stable_idx = df.loc[end_idx+1:][df['bug'] == 'Stable'].index.min()
    # 添加整个区间的索引
    keep_indices.update(range(prev_stable_idx, next_stable_idx + 1))

# 生成结果并重置索引
result_df = df.loc[sorted(keep_indices)].reset_index(drop=True)
# 输出结果(仅保留commit和bug列)
print(result_df[['commit', 'bug']])

逻辑说明

  1. 标记所有非Stable行,通过连续变化生成块ID,将连续的非Stable行归为同一组。
  2. 对每个块检查是否同时包含Increase和Decrease,筛选出有效块。
  3. 为每个有效块扩展范围:向前取最近的Stable行,向后取最近的Stable行,保留整个区间的所有行。

方案二:纯Python实现(无第三方依赖,适合小型数据集)

直接遍历列表,定位符合条件的区间:

# 原始数据集
data = [
    ('sha_1', 'Stable'),
    ('sha_2', 'Stable'),
    ('sha_3', 'Stable'),
    ('sha_4', 'Increase'),
    ('sha_5', 'Stable'),
    ('sha_6', 'Stable'),
    ('sha_7', 'Decrease'),
    ('sha_8', 'Stable'),
    ('sha_9', 'Decrease'),
    ('sha_10', 'Decrease'),
    ('sha_11', 'Increase'),
    ('sha_12', 'Stable'),
]

valid_ranges = []
current_start = None
current_types = set()

# 遍历数据集,定位有效非稳定区间
for idx, (commit, bug) in enumerate(data):
    if bug in ['Increase', 'Decrease']:
        if current_start is None:
            current_start = idx
        current_types.add(bug)
    else:
        if current_start is not None:
            # 检查当前区间是否包含两种类型
            if len(current_types) == 2:
                # 向前查找最近的Stable行
                prev_stable_idx = idx - 1
                while prev_stable_idx >= 0 and data[prev_stable_idx][1] != 'Stable':
                    prev_stable_idx -= 1
                # 向后的Stable行就是当前idx
                valid_ranges.append((prev_stable_idx, idx))
            # 重置当前区间状态
            current_start = None
            current_types = set()

# 收集结果
result = []
for start, end in valid_ranges:
    result.extend(data[start:end+1])

# 打印结果
for item in result:
    print(f"{item[0]} | {item[1]}")

逻辑说明

  1. 遍历过程中记录连续的非Stable区间,同时跟踪区间内的类型。
  2. 当遇到Stable行时,检查之前的非Stable区间是否包含两种类型,若是则扩展区间到前后的Stable行。
  3. 最后收集所有有效区间的行,输出结果。

内容的提问来源于stack exchange,提问作者Giammaria GIORDANO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:10:40