如何在Pandas中按组筛选符合跨度与包含关系的字符串行?
问题背景
我有如下样本DataFrame:
| identifier | span | matched_string | | -------- | -------------- | ------ | | occupation | [0,12] | general manager| | occupation | [0,7] | manager | | time schedule | [13,14] | "0-5" | | occupation | [0,12] | clerk |
需求说明
需按identifier分组,仅保留以下两类行:
- 同组内某
matched_string被其他字符串包含,且保留更长的那个字符串,同时两者的span存在重叠; - 同组内无包含关系的字符串行(如示例中的
time schedule行)。
最终目标DataFrame如下:
| identifier | span | matched_string | | -------- | -------------- | ------ | | occupation | [0,12] | general manager| | time schedule| [13,14] | "0-5" | | occupation | [0,12] | clerk |
现有代码
我尝试用以下代码处理,但嵌套循环过于繁琐:
ka = testa.groupby(["identifier"]) for name, group in ka: for index, row in group.iterrows(): for i in group['matched_string'].values: for j in group['matched_string'].values: for k in group['span_info'].values: for l in group['span_info'].values: if i != j and k != l: for m in list(range(k[0],k[1])): if m in list(range(l[0], l[1])) and i in j: max(i, j, key=len)
疑问
是否有更高效的实现方式?另外如何保留无包含关系的行?
解决方案
可以利用pandas的分组(groupby)结合自定义函数实现,避免多层嵌套循环,提升效率:
步骤1:定义分组处理函数
import pandas as pd def filter_group(group): group = group.reset_index(drop=True) keep = [True] * len(group) for idx in range(len(group)): current_str = group.loc[idx, 'matched_string'] current_span = group.loc[idx, 'span'] current_len = len(current_str) # 检查同组其他行是否满足"更长、包含当前字符串、span重叠" for other_idx in range(len(group)): if idx == other_idx: continue other_str = group.loc[other_idx, 'matched_string'] other_span = group.loc[other_idx, 'span'] other_len = len(other_str) # 判断span是否有重叠 spans_overlap = not (current_span[1] <= other_span[0] or other_span[1] <= current_span[0]) # 若存在更长的字符串包含当前且span重叠,则标记当前行不保留 if other_len > current_len and current_str in other_str and spans_overlap: keep[idx] = False break return group[keep]
步骤2:应用函数到分组数据
# 假设原DataFrame名为testa result = testa.groupby('identifier', group_keys=False).apply(filter_group)
逻辑说明
- 对每个分组内的每行,检查是否存在其他行满足字符串更长、包含当前字符串、span区间重叠,如果存在则丢弃当前行;
- 无包含关系的行(比如单一行的分组、同组内互相不包含的行)会自动保留;
- 相比原代码的多层嵌套,该方法减少了重复遍历,逻辑更清晰,效率更高。
内容的提问来源于stack exchange,提问作者julez8000
相关产品推荐
相关产品推荐

