You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

统计DataFrame每行中两指定短语在N词距内的匹配次数

需求:统计DataFrame每行中短语近距离匹配次数

需统计DataFrame每行字符串内,两个特定短语在指定词数距离内出现的次数,短语顺序无关。

示例说明

设短语X="black cat",Y="is my",指定词距为3:

  • 字符串The black cat is my black cat的匹配次数为2
  • 字符串The black cat by the window is my black cat的匹配次数为2
  • 字符串The black cat by the big window is my black cat的匹配次数为1

示例数据、问题代码与期望输出

示例数据

import pandas as pd

data = [['ABC123', 'test sentence here has these test words'], 
        ['ABC456', 'test sentence here contains these test words in test sentence form'], 
        ['ABC789', 'the third test sentence has no more additional test words']]
df = pd.DataFrame(data, columns=['Record ID', 'String'])
print(df)

输出的DataFrame结构:

Record IDString
ABC123test sentence here has these test words
ABC456test sentence here contains these test words in test sentence form
ABC789the third test sentence has no more additional test words

存在问题的代码

def phrase_finder(df, text_column, search_phrase, near_phrase, distance):
    results = 0
    for text in df[text_column]:
        for substring in text.split(search_phrase):
            words = substring.split()
            if len(words) <= distance + 1 and near_phrase in substring:
                results += 1
    return results if results else None

search_phrase = "test sentence"
near_phrase = "test words"
distance = 3

print(phrase_finder(df, 'String', search_phrase, near_phrase, distance))

该代码的核心问题:

  • 仅支持search_phrase在前、near_phrase在后的匹配,不满足顺序无关要求
  • 将所有行的结果累加,无法按行输出匹配次数
  • 匹配逻辑存在漏洞,无法精准计算两个短语间的词数距离

期望输出

IDNumber of Matches
ABC1231
ABC4562
ABC7890

解决方案

通过拆分单词、定位短语位置、计算词距的方式实现顺序无关的匹配统计:

import pandas as pd

def count_phrase_pairs(text, phrase1, phrase2, max_distance):
    # 拆分短语与文本为单词列表
    phrase1_words = phrase1.split()
    phrase2_words = phrase2.split()
    text_words = text.split()
    len_p1 = len(phrase1_words)
    len_p2 = len(phrase2_words)
    count = 0
    
    # 记录phrase1的所有起始索引
    p1_positions = []
    for i in range(len(text_words) - len_p1 + 1):
        if text_words[i:i+len_p1] == phrase1_words:
            p1_positions.append(i)
    
    # 记录phrase2的所有起始索引
    p2_positions = []
    for i in range(len(text_words) - len_p2 + 1):
        if text_words[i:i+len_p2] == phrase2_words:
            p2_positions.append(i)
    
    # 遍历短语位置对,计算词距并统计符合条件的次数
    for p1_start in p1_positions:
        p1_end = p1_start + len_p1 - 1
        for idx, p2_start in enumerate(p2_positions):
            p2_end = p2_start + len_p2 - 1
            # 计算两个短语之间的实际词数距离(排除短语自身单词)
            if p1_end < p2_start:
                distance = p2_start - p1_end - 1
            else:
                distance = p1_start - p2_end - 1
            if distance <= max_distance:
                count += 1
                # 移除已匹配的位置,避免重复计数
                del p2_positions[idx]
                break
    return count

# 将函数应用到DataFrame每行
search_phrase = "test sentence"
near_phrase = "test words"
distance = 3

df['Number of Matches'] = df['String'].apply(lambda x: count_phrase_pairs(x, search_phrase, near_phrase, distance))
print(df[['Record ID', 'Number of Matches']].rename(columns={'Record ID': 'ID'}))

代码说明

  1. 拆分与定位:将文本和目标短语拆分为单词列表,遍历文本找出每个短语的所有起始位置
  2. 词距计算:对每一对短语位置,计算二者之间的非短语单词数量,判断是否符合指定距离要求
  3. 去重计数:每匹配一对短语就移除已匹配的位置,避免同一对短语被重复统计
  4. 逐行处理:用apply函数对DataFrame每行单独计算匹配次数,输出按行统计的结果

运行后将输出符合期望的结果:

ID  Number of Matches
0  ABC123                  1
1  ABC456                  2
2  ABC789                  0

内容的提问来源于stack exchange,提问作者tshobe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:15:38