统计DataFrame每行中两指定短语在N词距内的匹配次数
需求:统计DataFrame每行中短语近距离匹配次数
需统计DataFrame每行字符串内,两个特定短语在指定词数距离内出现的次数,短语顺序无关。
示例说明
设短语X="black cat",Y="is my",指定词距为3:
- 字符串
The black cat is my black cat的匹配次数为2 - 字符串
The black cat by the window is my black cat的匹配次数为2 - 字符串
The black cat by the big window is my black cat的匹配次数为1
示例数据、问题代码与期望输出
示例数据
import pandas as pd data = [['ABC123', 'test sentence here has these test words'], ['ABC456', 'test sentence here contains these test words in test sentence form'], ['ABC789', 'the third test sentence has no more additional test words']] df = pd.DataFrame(data, columns=['Record ID', 'String']) print(df)
输出的DataFrame结构:
| Record ID | String |
|---|---|
| ABC123 | test sentence here has these test words |
| ABC456 | test sentence here contains these test words in test sentence form |
| ABC789 | the third test sentence has no more additional test words |
存在问题的代码
def phrase_finder(df, text_column, search_phrase, near_phrase, distance): results = 0 for text in df[text_column]: for substring in text.split(search_phrase): words = substring.split() if len(words) <= distance + 1 and near_phrase in substring: results += 1 return results if results else None search_phrase = "test sentence" near_phrase = "test words" distance = 3 print(phrase_finder(df, 'String', search_phrase, near_phrase, distance))
该代码的核心问题:
- 仅支持
search_phrase在前、near_phrase在后的匹配,不满足顺序无关要求 - 将所有行的结果累加,无法按行输出匹配次数
- 匹配逻辑存在漏洞,无法精准计算两个短语间的词数距离
期望输出
| ID | Number of Matches |
|---|---|
| ABC123 | 1 |
| ABC456 | 2 |
| ABC789 | 0 |
解决方案
通过拆分单词、定位短语位置、计算词距的方式实现顺序无关的匹配统计:
import pandas as pd def count_phrase_pairs(text, phrase1, phrase2, max_distance): # 拆分短语与文本为单词列表 phrase1_words = phrase1.split() phrase2_words = phrase2.split() text_words = text.split() len_p1 = len(phrase1_words) len_p2 = len(phrase2_words) count = 0 # 记录phrase1的所有起始索引 p1_positions = [] for i in range(len(text_words) - len_p1 + 1): if text_words[i:i+len_p1] == phrase1_words: p1_positions.append(i) # 记录phrase2的所有起始索引 p2_positions = [] for i in range(len(text_words) - len_p2 + 1): if text_words[i:i+len_p2] == phrase2_words: p2_positions.append(i) # 遍历短语位置对,计算词距并统计符合条件的次数 for p1_start in p1_positions: p1_end = p1_start + len_p1 - 1 for idx, p2_start in enumerate(p2_positions): p2_end = p2_start + len_p2 - 1 # 计算两个短语之间的实际词数距离(排除短语自身单词) if p1_end < p2_start: distance = p2_start - p1_end - 1 else: distance = p1_start - p2_end - 1 if distance <= max_distance: count += 1 # 移除已匹配的位置,避免重复计数 del p2_positions[idx] break return count # 将函数应用到DataFrame每行 search_phrase = "test sentence" near_phrase = "test words" distance = 3 df['Number of Matches'] = df['String'].apply(lambda x: count_phrase_pairs(x, search_phrase, near_phrase, distance)) print(df[['Record ID', 'Number of Matches']].rename(columns={'Record ID': 'ID'}))
代码说明
- 拆分与定位:将文本和目标短语拆分为单词列表,遍历文本找出每个短语的所有起始位置
- 词距计算:对每一对短语位置,计算二者之间的非短语单词数量,判断是否符合指定距离要求
- 去重计数:每匹配一对短语就移除已匹配的位置,避免同一对短语被重复统计
- 逐行处理:用
apply函数对DataFrame每行单独计算匹配次数,输出按行统计的结果
运行后将输出符合期望的结果:
ID Number of Matches 0 ABC123 1 1 ABC456 2 2 ABC789 0
内容的提问来源于stack exchange,提问作者tshobe
相关产品推荐
相关产品推荐

