如何实现模糊字符串匹配?Python英文词典输入纠错需求
英文词典模糊匹配实现方案
用户需求与现有代码
我正在开发一个英文词典,现有实现代码如下:
import json file = open("data.json", "r", encoding= "utf-8") #data.json contains some words and meanings in English as a dict. dictionary = json.loads(file.read()) word = input("Enter Word: ") print(dictionary[word])
希望在用户输入单词有误时,显示字典中相近的键(例如输入runn时,提示“你是不是指‘run’”),询问是否有可用的模糊字符串匹配函数。
实现方案
1. 用Python标准库difflib(推荐)
Python自带的difflib.get_close_matches函数可以直接从候选键中找出与输入最相似的项,无需额外安装依赖,适合大多数场景。
修改后的完整代码:
import json import difflib # 安全读取字典文件 with open("data.json", "r", encoding="utf-8") as file: dictionary = json.load(file) input_word = input("Enter Word: ").lower() if input_word in dictionary: print(dictionary[input_word]) else: # 获取最相似的1个匹配项,cutoff为相似度阈值(0-1,值越高要求越像) close_matches = difflib.get_close_matches(input_word, dictionary.keys(), n=1, cutoff=0.6) if close_matches: print(f"你是不是指‘{close_matches[0]}’?") else: print("未找到该单词,也没有相近的匹配项。")
with语句自动管理文件关闭,比直接open更安全。- 统一转为小写,避免大小写差异导致的匹配失败。
cutoff参数可根据需求调整,比如调至0.7会要求更高的相似度。
2. 自定义编辑距离算法(灵活定制)
如果需要自定义相似度规则,可以实现莱文斯坦距离(衡量两个字符串的编辑差异次数),手动找出差异最小的单词:
import json def levenshtein_distance(s1, s2): if len(s1) < len(s2): return levenshtein_distance(s2, s1) if len(s2) == 0: return len(s1) previous_row = range(len(s2) + 1) for i, c1 in enumerate(s1): current_row = [i + 1] for j, c2 in enumerate(s2): insert_cost = previous_row[j + 1] + 1 delete_cost = current_row[j] + 1 substitute_cost = previous_row[j] + (c1 != c2) current_row.append(min(insert_cost, delete_cost, substitute_cost)) previous_row = current_row return previous_row[-1] # 读取字典数据 with open("data.json", "r", encoding="utf-8") as file: dictionary = json.load(file) input_word = input("Enter Word: ").lower() if input_word in dictionary: print(dictionary[input_word]) else: # 找出编辑距离最小的单词 closest_word = min(dictionary.keys(), key=lambda x: levenshtein_distance(input_word, x.lower())) print(f"你是不是指‘{closest_word}’?")
这种方式适合需要特殊匹配逻辑的场景,但代码量比使用difflib更大。
内容的提问来源于stack exchange,提问作者RestartGame RG
相关产品推荐
相关产品推荐

