You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按组筛选符合跨度与包含关系的字符串行?

问题背景

我有如下样本DataFrame:

| identifier    | span           | matched_string |
| --------      | -------------- | ------         | 
| occupation    | [0,12]         | general manager| 
| occupation    | [0,7]          | manager        |
| time schedule | [13,14]        | "0-5"          |
| occupation    | [0,12]         | clerk          |
需求说明

需按identifier分组,仅保留以下两类行:

  • 同组内某matched_string被其他字符串包含,且保留更长的那个字符串,同时两者的span存在重叠;
  • 同组内无包含关系的字符串行(如示例中的time schedule行)。

最终目标DataFrame如下:

| identifier   | span           | matched_string |
| --------     | -------------- | ------         |
| occupation   | [0,12]         | general manager| 
| time schedule| [13,14]        | "0-5"          |
| occupation   | [0,12]         | clerk          |
现有代码

我尝试用以下代码处理,但嵌套循环过于繁琐:

ka = testa.groupby(["identifier"])
    
for name, group in ka:
      for index, row in group.iterrows(): 
            for i in group['matched_string'].values:
                for j in group['matched_string'].values:
                    for k in group['span_info'].values:
                        for l in group['span_info'].values:
                            if i != j and k != l:
                                for m in list(range(k[0],k[1])):
                                    if m in list(range(l[0], l[1])) and i in j:
                                       max(i, j, key=len)
疑问

是否有更高效的实现方式?另外如何保留无包含关系的行?

解决方案

可以利用pandas的分组(groupby)结合自定义函数实现,避免多层嵌套循环,提升效率:

步骤1:定义分组处理函数

import pandas as pd

def filter_group(group):
    group = group.reset_index(drop=True)
    keep = [True] * len(group)
    
    for idx in range(len(group)):
        current_str = group.loc[idx, 'matched_string']
        current_span = group.loc[idx, 'span']
        current_len = len(current_str)
        
        # 检查同组其他行是否满足"更长、包含当前字符串、span重叠"
        for other_idx in range(len(group)):
            if idx == other_idx:
                continue
            other_str = group.loc[other_idx, 'matched_string']
            other_span = group.loc[other_idx, 'span']
            other_len = len(other_str)
            
            # 判断span是否有重叠
            spans_overlap = not (current_span[1] <= other_span[0] or other_span[1] <= current_span[0])
            # 若存在更长的字符串包含当前且span重叠,则标记当前行不保留
            if other_len > current_len and current_str in other_str and spans_overlap:
                keep[idx] = False
                break
    
    return group[keep]

步骤2:应用函数到分组数据

# 假设原DataFrame名为testa
result = testa.groupby('identifier', group_keys=False).apply(filter_group)

逻辑说明

  • 对每个分组内的每行,检查是否存在其他行满足字符串更长、包含当前字符串、span区间重叠,如果存在则丢弃当前行;
  • 无包含关系的行(比如单一行的分组、同组内互相不包含的行)会自动保留;
  • 相比原代码的多层嵌套,该方法减少了重复遍历,逻辑更清晰,效率更高。

内容的提问来源于stack exchange,提问作者julez8000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:10:29