如何在Python列表中提取特定字符B后直到遇到B/O的所有索引
BIO序列索引分组实现方案
遍历是该场景下时间复杂度最优的实现方式(O(n),仅需遍历一次列表),我们可以通过简单的状态标记实现非常简洁的代码:
推荐实现(性能最优,可读性高)
def group_bio_indices(data): result = [] current_group = None for idx, tag in enumerate(data): if tag == 'B': # 遇到B开启新分组 current_group = [idx] result.append(current_group) elif tag == 'I' and current_group is not None: # 处于有效分组中遇到I,追加到当前分组 current_group.append(idx) else: # 遇到O,关闭当前分组 current_group = None return result
测试验证
- 测试用例1:
data = ['O', 'O', 'B', 'I', 'I', 'B', 'I', 'O', 'B', 'I'] print(group_bio_indices(data)) # 输出:[[2, 3, 4], [5, 6], [8, 9]]
- 测试用例2:
data = ['B', 'I', 'I', 'O', 'O', 'B', 'I', 'B', 'I', 'I', 'O'] print(group_bio_indices(data)) # 输出:[[0, 1, 2], [5, 6], [7, 8, 9]]
可选短写法(基于itertools.groupby)
如果偏好函数式写法,也可以用itertools.groupby实现,性能与普通遍历持平:
from itertools import groupby def group_bio_indices(data): result = [] for is_valid, group in groupby(enumerate(data), key=lambda x: x[1] != 'O'): if not is_valid: continue idx_list = [i for i, _ in group] sub_group = [] for i in idx_list: if data[i] == 'B': if sub_group: result.append(sub_group) sub_group = [i] else: sub_group.append(i) if sub_group: result.append(sub_group) return result
内容的提问来源于stack exchange,提问作者Muhammad Fhadli
相关产品推荐
相关产品推荐

