如何从字符串列表匹配目标值或前3个相似字符串,解决difflib.get_close_matches匹配错误问题
带位置占位符的字符串匹配问题解决方案
问题根源
difflib.get_close_matches的匹配逻辑基于整体字符串的编辑距离,不会识别输入中的_作为单字符占位符,也不会优先匹配固定位置的字符,同时你设置的0.2的cutoff阈值过低,会拉出大量完全不相关的结果。
实现方案
核心思路是先通过正则匹配固定位置的字符,再按同位置匹配字符数排序输出结果,完全符合按给定字符位置匹配的需求:
代码实现
import re def match_pokemon(hint, candidates, n=3): # 1. 先筛选和提示长度相同的候选,排除长度不匹配的无效结果 hint_len = len(hint) same_len_candidates = [p for p in candidates if len(p) == hint_len] if not same_len_candidates: return [] # 2. 把提示的下划线转成正则通配符,锚定首尾保证完整匹配 pattern_str = '^' + re.escape(hint).replace(r'\_', '.') + '$' pattern = re.compile(pattern_str, re.IGNORECASE) # 忽略大小写可按需关闭 # 3. 先找完全匹配位置规则的结果 exact_matches = [p for p in same_len_candidates if pattern.fullmatch(p)] # 4. 剩余名额用同位置匹配字符数排序补全 remaining = n - len(exact_matches) if remaining <= 0: return exact_matches[:n] # 排除已经匹配到的结果,计算剩余候选和hint的同位置匹配度 rest_candidates = [p for p in same_len_candidates if p not in exact_matches] # 按匹配字符数从高到低排序 rest_candidates.sort(key=lambda x: sum(1 for i in range(hint_len) if x[i].lower() == hint[i].lower() and hint[i] != '_'), reverse=True) # 合并结果 return exact_matches + rest_candidates[:remaining] # 调用方式替换原来的get_close_matches即可 res = match_pokemon(hint, pokemons, n=3) if not res: print("No Similar Names Found Unfortunately") else: var = "\n".join(res) print(f"**Did you Mean...**\n`{var}`")
效果验证
针对你给出的示例hint="_ar__n_e",该函数会先过滤出所有长度为8的宝可梦名称,再通过正则^.ar..n.e$匹配,直接就能命中mareanie作为首个结果,后续补充的结果也是同位置匹配字符最多的相关名称,不会出现无关结果。
内容的提问来源于stack exchange,提问作者Prabsimar
相关产品推荐
相关产品推荐

