如何在字符串中检索目标短语及其近似变体并提取所有匹配结果
需求说明
现有如下待检索字符串:
source_str = 'machine learning ml is a type of artificial intelligence ai that allows software applications to become more accurate at predicting outcomes without being explicitly programmed to do so machine12 learning algorithms use historical data as input to predict new output values machines learning is good'
待匹配的目标标签为:
target_tag = 'machine learning'
需要检索目标标签及其所有近似匹配项,包含示例中的machine learning、machine12 learning、machines learning,最终输出匹配结果列表。
已有方案的缺陷
之前尝试的正则匹配只能覆盖特定规则的变体:
- 匹配中间为空格/数字的规则
r"(machine[\s0-9]+learning)",只能匹配到machine learning、machine12 learning,漏了machines learning - 匹配中间为空格/字母的规则
r"(machine[\sA-Za-z]+learning)",只能匹配到machine learning、machines learning,漏了machine12 learning
这类硬编码规则的通用性极差,无法适配更多未知的近似场景。
通用实现方案
推荐使用滑动窗口+模糊相似度匹配的方案,不需要预先定义规则,可适配各类近似变体,实现步骤如下:
首先安装依赖库:
pip install fuzzywuzzy python-Levenshtein nltk
完整实现代码:
import nltk from fuzzywuzzy import fuzz # 首次运行需要下载分词依赖包,后续可注释该行 nltk.download('punkt') source_str = 'machine learning ml is a type of artificial intelligence ai that allows software applications to become more accurate at predicting outcomes without being explicitly programmed to do so machine12 learning algorithms use historical data as input to predict new output values machines learning is good' target_tag = 'machine learning' # 相似度阈值,可根据匹配精度需求灵活调整,分值范围0-100 similarity_threshold = 75 # 目标标签分词,获取标签的单词数量 tag_tokens = nltk.word_tokenize(target_tag.lower()) tag_word_count = len(tag_tokens) # 源文本分词 source_tokens = nltk.word_tokenize(source_str) match_results = [] # 支持窗口比标签多1个单词,适配中间插入短字符/数字的场景 for window_size in range(tag_word_count, tag_word_count + 2): # 遍历所有滑动窗口 for idx in range(len(source_tokens) - window_size + 1): current_window_str = ' '.join(source_tokens[idx:idx+window_size]) # 计算当前窗口和目标标签的编辑距离相似度 if fuzz.ratio(current_window_str.lower(), target_tag.lower()) >= similarity_threshold: # 去重后加入结果列表 if current_window_str not in match_results: match_results.append(current_window_str) print(match_results) # 输出结果:['machine learning', 'machine12 learning', 'machines learning']
方案优势
- 适配性强:无需提前定义正则规则,可自动覆盖单词变形、中间插入数字/符号、大小写差异等各类近似场景
- 灵活可控:可通过调整
similarity_threshold阈值平衡召回率和准确率,要求严格就调高阈值,需要多召回近似结果就调低阈值
如果需要适配语义级别的近似匹配(比如同义词替换、单词顺序颠倒),把模糊相似度替换为词向量余弦相似度校验即可。
内容的提问来源于stack exchange,提问作者Wiliam
相关产品推荐
相关产品推荐

