You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python列表中提取特定字符B后直到遇到B/O的所有索引

BIO序列索引分组实现方案

遍历是该场景下时间复杂度最优的实现方式(O(n),仅需遍历一次列表),我们可以通过简单的状态标记实现非常简洁的代码:

推荐实现(性能最优,可读性高)

def group_bio_indices(data):
    result = []
    current_group = None
    for idx, tag in enumerate(data):
        if tag == 'B':
            # 遇到B开启新分组
            current_group = [idx]
            result.append(current_group)
        elif tag == 'I' and current_group is not None:
            # 处于有效分组中遇到I,追加到当前分组
            current_group.append(idx)
        else:
            # 遇到O,关闭当前分组
            current_group = None
    return result

测试验证

  • 测试用例1:
data = ['O', 'O', 'B', 'I', 'I', 'B', 'I', 'O', 'B', 'I']
print(group_bio_indices(data)) 
# 输出:[[2, 3, 4], [5, 6], [8, 9]]
  • 测试用例2:
data = ['B', 'I', 'I', 'O', 'O', 'B', 'I', 'B', 'I', 'I', 'O']
print(group_bio_indices(data))
# 输出:[[0, 1, 2], [5, 6], [7, 8, 9]]

可选短写法(基于itertools.groupby)

如果偏好函数式写法,也可以用itertools.groupby实现,性能与普通遍历持平:

from itertools import groupby

def group_bio_indices(data):
    result = []
    for is_valid, group in groupby(enumerate(data), key=lambda x: x[1] != 'O'):
        if not is_valid:
            continue
        idx_list = [i for i, _ in group]
        sub_group = []
        for i in idx_list:
            if data[i] == 'B':
                if sub_group:
                    result.append(sub_group)
                sub_group = [i]
            else:
                sub_group.append(i)
        if sub_group:
            result.append(sub_group)
    return result

内容的提问来源于stack exchange,提问作者Muhammad Fhadli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 19:15:08