Python(Pandas)统计两词相距N词内的出现次数(含正则/非正则方案)
解决方案:统计指定单词在N词范围内的共现次数
一、修正你的错误代码
你的代码出错是因为循环遍历了DataFrame的列名而非行数据。以下是修正后的版本:
import re import pandas as pd # 假设你的nearwordDF已经预处理完成 output = [] # 使用iterrows()遍历每一行 for idx, row in nearwordDF.iterrows(): text = row['text'] # 匹配"my"和"money"在2个词范围内的共现(包括相邻) regex = r'(?=\b(?:my\s+(?:\w+\s+){0,2}money|money\s+(?:\w+\s+){0,2}my)\b)' matches = re.findall(regex, text, flags=re.IGNORECASE) output.append([row['id'], len(matches)]) df = pd.DataFrame(output, columns=['Record ID', 'Occurrences']) print(df)
修正点说明:
- 用
iterrows()遍历DataFrame的每一行,而非直接遍历列名 - 改用零宽正向预查(
(?=...))的正则,避免遗漏重叠的共现情况 - 添加
re.IGNORECASE标志实现大小写不敏感匹配
二、通用正则枚举函数
以下是可复用的正则函数,支持自定义目标单词和间隔词数N:
import re def count_near_words_regex(text, word1, word2, max_words_between): # 转义特殊字符,避免正则语法冲突 w1 = re.escape(word1) w2 = re.escape(word2) # 构建正则模式:匹配word1后接0到max_words_between个词再到word2,反之亦然 # 使用零宽预查确保不遗漏重叠匹配 pattern = fr''' (?=\b(?: {w1}\W+(?:\w+\W+)?{{0,{max_words_between}}}{w2}| {w2}\W+(?:\w+\W+)?{{0,{max_words_between}}}{w1} )\b) ''' # 查找所有匹配项,返回数量 matches = re.findall(pattern, text, flags=re.IGNORECASE | re.VERBOSE) return len(matches) # 在Pandas中应用该函数 df['Occurrences Identified'] = df['String'].apply( lambda x: count_near_words_regex(x, 'sentence', 'the', 3) )
三、非正则简化解决方案
如果觉得正则难以维护,可采用拆分单词+遍历配对的方式,逻辑更直观:
def count_near_words_simple(text, word1, word2, max_words_between): # 拆分字符串为单词列表(大小写统一为小写) words = text.lower().split() # 收集所有目标单词的位置和内容 target_positions = [] for idx, word in enumerate(words): if word in (word1.lower(), word2.lower()): target_positions.append((idx, word)) count = 0 # 遍历所有两两组合,统计符合条件的异词配对 for i in range(len(target_positions)): pos_i, word_i = target_positions[i] for j in range(i+1, len(target_positions)): pos_j, word_j = target_positions[j] # 两个单词不同,且间隔不超过max_words_between个词 if word_i != word_j and abs(pos_i - pos_j) <= max_words_between + 1: count += 1 return count # 在Pandas中应用该函数 df['Occurrences Identified'] = df['String'].apply( lambda x: count_near_words_simple(x, 'sentence', 'the', 3) )
逻辑说明:
- 将文本拆分为单词列表,统一大小写
- 记录所有目标单词的位置
- 检查每一对不同单词的位置差:若位置差≤
max_words_between+1(即中间最多有max_words_between个词),则计数加1 - 避免重复计数(只遍历i<j的组合)
测试验证
针对你的示例输入:
data = [['ABC123', 'This is the first example sentence the end of sentence one'], ['ABC456', 'This is the second example sentence one more sentence to come'], ['ABC789', 'There are no more example sentences']] df = pd.DataFrame(data, columns=['Record ID', 'String'])
两种方法都会输出符合预期的结果:
| Record ID | Occurrences Identified |
|---|---|
| ABC123 | 3 |
| ABC456 | 1 |
| ABC789 | 0 |
内容的提问来源于stack exchange,提问作者tshobe
相关产品推荐
相关产品推荐

