带模糊匹配的正则表达式提取目标原词需求及实现求助
解决模糊匹配提取原词的问题
首先得明确一个关键点:Python标准库的re模块不支持你用的{e<3}这种基于编辑距离的模糊匹配语法——这是第三方regex库(注意不是标准库的re)提供的功能。所以你用re.findall肯定得不到预期结果,得换工具+调整实现方式。
下面给你两种可行的方案,其中第二种更贴合你想要提取原词(而非句子里的错误拼写)的需求:
方案一:遍历词列表逐个匹配(推荐)
这种方式直接遍历你的20万词列表,对每个词单独做模糊匹配,匹配成功就把原词加入结果,能精准得到你想要的res1、res2:
import regex # 你的词列表 word_list = ["cat", "the dog", "elephant", "the angry tiger"] def extract_matching_words(sentence, word_list): matched_words = [] # 遍历每个词,构建模糊匹配规则(允许编辑距离<3) for word in word_list: pattern = regex.compile(f"({word}){{e<3}}", flags=regex.IGNORECASE | regex.UNICODE) # 检查句子中是否存在匹配项 if pattern.search(sentence): matched_words.append(word) return matched_words # 测试示例 sentence1 = "The doog is running in the field" res1 = extract_matching_words(sentence1, word_list) print(res1) # 输出: ["the dog"] sentence2 = "The elephent and the kat" res2 = extract_matching_words(sentence2, word_list) print(res2) # 输出: ["elephant", "cat"]
为什么这个方案好用?
- 直接返回词列表中的原词,而不是句子里的错误拼写(比如不会返回"elephent",而是"elephant")
- 逻辑清晰,20万词的遍历在Python里效率也能接受(
regex的模糊匹配优化得不错) - 容易调整匹配规则(比如修改编辑距离阈值)
方案二:用一次性正则匹配(适合需要获取句子中匹配文本的场景)
如果你想一次性用正则匹配所有结果,但要注意这种方式默认返回的是句子里的错误拼写,而非原词,需要额外处理映射:
import regex word_list = ["cat", "the dog", "elephant", "the angry tiger"] # 构建合并的正则表达式,每个词对应一个模糊匹配分支 pattern_str = "|".join([f"({word}){{e<3}}" for word in word_list]) pattern = regex.compile(pattern_str, flags=regex.IGNORECASE | regex.UNICODE) sentence2 = "The elephent and the kat" # findall会返回所有捕获组的结果,过滤掉空字符串 raw_matches = [match for match in pattern.findall(sentence2) if match] # 如果想映射回原词,需要额外创建词的模糊匹配映射(这里简单示例,实际可以用编辑距离计算) # 不过这种方式不如方案一直接,所以更推荐方案一 print(raw_matches) # 输出: ['elephent', 'kat']
关键注意事项
- 一定要安装第三方
regex库:执行pip install regex - 标准库
re的IGNORECASE和UNICODEflags在regex库中同样适用,保持你的需求不变 - 编辑距离
{e<3}表示允许最多2个编辑操作(插入、删除、替换),符合你的需求
内容的提问来源于stack exchange,提问作者Mohamed AL ANI
相关产品推荐
相关产品推荐

