Pandas DataFrame中基于正则提取关键词对的最小词距问题
解决Pandas中关键词对匹配不全及最小词距计算问题
我看了你的代码和问题描述,发现主要有两个核心问题导致匹配结果不全:
原函数的问题分析
- 结果扁平化,未按原字符串分组:你的函数把所有字符串的匹配结果合并成了一个列表,导致你无法看到每个单独字符串的匹配情况,误以为某些结果缺失。比如第一行第一个字符串里的
ghi jkl其实被匹配到了,但和其他字符串的结果混在一起,你可能没注意到。 - 正则仅匹配词对的正向顺序:你的正则
word1.*?word2只能匹配word1在word2之前的情况(虽然你的示例里没有反向的情况,但如果以后遇到会漏匹配),而且没有使用单词边界,可能会出现部分匹配的问题(比如把abcd当成abc匹配)。
修正方案:按字符串分组匹配,支持有序/无序词对
首先,我们需要调整函数逻辑,为每个输入字符串单独处理匹配,同时确保匹配完整单词。如果需要计算最小词距,我们还可以直接在函数中计算词之间的单词距离(而不是字符距离)。
方案1:获取所有匹配的词对片段(按字符串分组)
这个方案会返回每个字符串中所有匹配的关键词对片段,并且按原字符串顺序分组:
import pandas as pd import numpy as np import re # 定义关键词对列表 ListB = [['abc','def'],['ghi','jkl'],['mno','pqr']] # 构造测试数据 data = pd.DataFrame(np.array([['1', '2', ['random string to be searched abc def ghi jkl','random string to be searched abc','abc random string to be searched def']], ['4', '5', ['random string to be searched ghi jkl','random string to be searched',' mno random string to be searched pqr']], ['7', '8', ['abc random string to be searched def','random string to be searched mno pqr','random string to be searched']]]), columns=['a', 'b', 'list_of_strings_to_search']) def find_matching_pairs(text_list, word_pairs): """ 为每个字符串找出所有匹配的关键词对片段,返回按原字符串分组的结果 """ grouped_results = [] for text in text_list: string_matches = [] for pair in word_pairs: word1, word2 = pair # 使用单词边界\b确保匹配完整单词,同时支持正向和反向顺序(如果不需要反向,移除|后的部分) pattern = re.compile(rf'\b(?:{re.escape(word1)}.*?{re.escape(word2)}|{re.escape(word2)}.*?{re.escape(word1)})\b', re.DOTALL) # 找到当前字符串中所有匹配的片段 matches = pattern.findall(text) string_matches.extend(matches) grouped_results.append(string_matches) return grouped_results # 应用函数到DataFrame data['matched_pairs'] = data['list_of_strings_to_search'].apply(find_matching_pairs, word_pairs=ListB) # 查看结果 print(data['matched_pairs'])
运行后,第一行的matched_pairs会是:
[ ['abc def', 'ghi jkl'], # 第一个字符串的匹配结果 [], # 第二个字符串只有abc,无匹配词对 ['abc random string to be searched def'] # 第三个字符串的匹配结果 ]
方案2:直接计算每个词对的最小词距
如果你的最终目标是计算词对间的最小词距(以单词数量为单位),可以使用下面的函数,它会直接返回每个字符串中存在的词对及其最小距离:
def find_min_word_distances(text_list, word_pairs): grouped_distances = [] for text in text_list: # 按空格分割单词(如果有标点,建议先清理,比如用re.sub(r'[^\w\s]', '', text)) words = text.strip().split() string_distances = [] for pair in word_pairs: word1, word2 = pair # 找到两个词在单词列表中的所有索引 word1_indices = [i for i, w in enumerate(words) if w == word1] word2_indices = [i for i, w in enumerate(words) if w == word2] if not word1_indices or not word2_indices: continue # 两个词不同时存在,跳过 # 计算所有可能的距离,取最小值 min_distance = min(abs(i - j) for i in word1_indices for j in word2_indices) string_distances.append({ 'word_pair': pair, 'min_word_distance': min_distance }) grouped_distances.append(string_distances) return grouped_distances # 应用函数到DataFrame data['min_word_distances'] = data['list_of_strings_to_search'].apply(find_min_word_distances, word_pairs=ListB) # 查看结果 print(data['min_word_distances'])
比如第一行第一个字符串的结果会是:
[ {'word_pair': ['abc', 'def'], 'min_word_distance': 1}, {'word_pair': ['ghi', 'jkl'], 'min_word_distance': 1} ]
关键改进点
- 使用
re.escape()处理关键词,避免关键词中的特殊字符破坏正则表达式 - 加入单词边界
\b确保匹配完整单词,避免部分匹配 - 按原字符串分组返回结果,清晰展示每个字符串的匹配情况
- 直接计算词距的方案更高效,不需要依赖正则匹配,结果更准确
内容的提问来源于stack exchange,提问作者DJW001
相关产品推荐
相关产品推荐

