You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现模糊字符串匹配?Python英文词典输入纠错需求

英文词典模糊匹配实现方案

用户需求与现有代码

我正在开发一个英文词典,现有实现代码如下:

import json
file = open("data.json", "r", encoding= "utf-8")

#data.json contains some words and meanings in English as a dict.

dictionary = json.loads(file.read())

word = input("Enter Word: ")

print(dictionary[word])

希望在用户输入单词有误时,显示字典中相近的键(例如输入runn时,提示“你是不是指‘run’”),询问是否有可用的模糊字符串匹配函数。

实现方案

1. 用Python标准库difflib(推荐)

Python自带的difflib.get_close_matches函数可以直接从候选键中找出与输入最相似的项,无需额外安装依赖,适合大多数场景。

修改后的完整代码:

import json
import difflib

# 安全读取字典文件
with open("data.json", "r", encoding="utf-8") as file:
    dictionary = json.load(file)

input_word = input("Enter Word: ").lower()

if input_word in dictionary:
    print(dictionary[input_word])
else:
    # 获取最相似的1个匹配项,cutoff为相似度阈值(0-1,值越高要求越像)
    close_matches = difflib.get_close_matches(input_word, dictionary.keys(), n=1, cutoff=0.6)
    if close_matches:
        print(f"你是不是指‘{close_matches[0]}’?")
    else:
        print("未找到该单词,也没有相近的匹配项。")
  • with语句自动管理文件关闭,比直接open更安全。
  • 统一转为小写,避免大小写差异导致的匹配失败。
  • cutoff参数可根据需求调整,比如调至0.7会要求更高的相似度。

2. 自定义编辑距离算法(灵活定制)

如果需要自定义相似度规则,可以实现莱文斯坦距离(衡量两个字符串的编辑差异次数),手动找出差异最小的单词:

import json

def levenshtein_distance(s1, s2):
    if len(s1) < len(s2):
        return levenshtein_distance(s2, s1)
    if len(s2) == 0:
        return len(s1)
    previous_row = range(len(s2) + 1)
    for i, c1 in enumerate(s1):
        current_row = [i + 1]
        for j, c2 in enumerate(s2):
            insert_cost = previous_row[j + 1] + 1
            delete_cost = current_row[j] + 1
            substitute_cost = previous_row[j] + (c1 != c2)
            current_row.append(min(insert_cost, delete_cost, substitute_cost))
        previous_row = current_row
    return previous_row[-1]

# 读取字典数据
with open("data.json", "r", encoding="utf-8") as file:
    dictionary = json.load(file)

input_word = input("Enter Word: ").lower()

if input_word in dictionary:
    print(dictionary[input_word])
else:
    # 找出编辑距离最小的单词
    closest_word = min(dictionary.keys(), key=lambda x: levenshtein_distance(input_word, x.lower()))
    print(f"你是不是指‘{closest_word}’?")

这种方式适合需要特殊匹配逻辑的场景,但代码量比使用difflib更大。

内容的提问来源于stack exchange,提问作者RestartGame RG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 05:42:25