You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python(Pandas)统计两词相距N词内的出现次数(含正则/非正则方案)

解决方案:统计指定单词在N词范围内的共现次数

一、修正你的错误代码

你的代码出错是因为循环遍历了DataFrame的列名而非行数据。以下是修正后的版本:

import re
import pandas as pd

# 假设你的nearwordDF已经预处理完成
output = []
# 使用iterrows()遍历每一行
for idx, row in nearwordDF.iterrows():
    text = row['text']
    # 匹配"my"和"money"在2个词范围内的共现(包括相邻)
    regex = r'(?=\b(?:my\s+(?:\w+\s+){0,2}money|money\s+(?:\w+\s+){0,2}my)\b)'
    matches = re.findall(regex, text, flags=re.IGNORECASE)
    output.append([row['id'], len(matches)])

df = pd.DataFrame(output, columns=['Record ID', 'Occurrences'])
print(df)

修正点说明:

  • 用iterrows()遍历DataFrame的每一行,而非直接遍历列名
  • 改用零宽正向预查((?=...))的正则,避免遗漏重叠的共现情况
  • 添加re.IGNORECASE标志实现大小写不敏感匹配

二、通用正则枚举函数

以下是可复用的正则函数,支持自定义目标单词和间隔词数N:

import re

def count_near_words_regex(text, word1, word2, max_words_between):
    # 转义特殊字符,避免正则语法冲突
    w1 = re.escape(word1)
    w2 = re.escape(word2)
    # 构建正则模式:匹配word1后接0到max_words_between个词再到word2,反之亦然
    # 使用零宽预查确保不遗漏重叠匹配
    pattern = fr'''
        (?=\b(?:
            {w1}\W+(?:\w+\W+)?{{0,{max_words_between}}}{w2}|
            {w2}\W+(?:\w+\W+)?{{0,{max_words_between}}}{w1}
        )\b)
    '''
    # 查找所有匹配项,返回数量
    matches = re.findall(pattern, text, flags=re.IGNORECASE | re.VERBOSE)
    return len(matches)

# 在Pandas中应用该函数
df['Occurrences Identified'] = df['String'].apply(
    lambda x: count_near_words_regex(x, 'sentence', 'the', 3)
)

三、非正则简化解决方案

如果觉得正则难以维护,可采用拆分单词+遍历配对的方式,逻辑更直观:

def count_near_words_simple(text, word1, word2, max_words_between):
    # 拆分字符串为单词列表(大小写统一为小写)
    words = text.lower().split()
    # 收集所有目标单词的位置和内容
    target_positions = []
    for idx, word in enumerate(words):
        if word in (word1.lower(), word2.lower()):
            target_positions.append((idx, word))
    
    count = 0
    # 遍历所有两两组合,统计符合条件的异词配对
    for i in range(len(target_positions)):
        pos_i, word_i = target_positions[i]
        for j in range(i+1, len(target_positions)):
            pos_j, word_j = target_positions[j]
            # 两个单词不同,且间隔不超过max_words_between个词
            if word_i != word_j and abs(pos_i - pos_j) <= max_words_between + 1:
                count += 1
    return count

# 在Pandas中应用该函数
df['Occurrences Identified'] = df['String'].apply(
    lambda x: count_near_words_simple(x, 'sentence', 'the', 3)
)

逻辑说明:

  1. 将文本拆分为单词列表,统一大小写
  2. 记录所有目标单词的位置
  3. 检查每一对不同单词的位置差:若位置差≤max_words_between+1(即中间最多有max_words_between个词),则计数加1
  4. 避免重复计数(只遍历i<j的组合)

测试验证

针对你的示例输入:

data = [['ABC123', 'This is the first example sentence the end of sentence one'], 
        ['ABC456', 'This is the second example sentence one more sentence to come'], 
        ['ABC789', 'There are no more example sentences']]
df = pd.DataFrame(data, columns=['Record ID', 'String'])

两种方法都会输出符合预期的结果:

Record IDOccurrences Identified
ABC1233
ABC4561
ABC7890

内容的提问来源于stack exchange,提问作者tshobe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 14:05:13