如何在Fuzzywuzzy字符串匹配项目中添加分数阈值过滤低匹配结果?
解决FuzzyWuzzy中
score_cutoff参数不生效及添加匹配分数阈值的问题 我来帮你搞定这个匹配分数阈值的需求!你遇到process.extract不接受score_cutoff参数的问题,多半是因为使用的fuzzywuzzy版本比较旧——这个参数是在后续版本中新增的特性。下面我给你两种解决方案,不管是升级版本还是手动过滤,都能实现只保留高分匹配结果的目标。
方案一:升级FuzzyWuzzy并使用score_cutoff参数
首先先把fuzzywuzzy升级到最新版本,建议同时安装python-Levenshtein来大幅提升匹配速度:
pip install --upgrade fuzzywuzzy python-Levenshtein
接下来修改你的代码,添加score_cutoff参数(比如设为80,你可以根据需求调整),同时修正原代码里的一些小问题(比如limit参数和函数的num_match不统一、列名生成逻辑优化等):
from fuzzywuzzy import process import pandas as pd import os def StringMatch(master, testfile, num_match=3, score_threshold=80): master_names = master.iloc[:,3] test_names = testfile.iloc[:,0] # 加入score_cutoff参数,只保留分数高于阈值的结果 fhp_new = [process.extract(x, master_names, limit=num_match, score_cutoff=score_threshold) for x in test_names] # 生成列名,更简洁的方式 match_columns = [f"Match{i}" for i in range(1, num_match+1)] aggregated_matches = pd.DataFrame(fhp_new, columns=match_columns) d = {} for col in match_columns: # 处理可能的空匹配(如果没有达到阈值的结果,填充为None) d[col] = [y[0] if y else None for y in aggregated_matches[col]] d["test_original"] = test_names.values # 完美匹配判断也要考虑空值的情况 d["perfect match"] = d["Match1"] == d["test_original"] out = pd.DataFrame(data=d) out.to_csv(f"{outFile}.csv", index=False) return out print("starting...") master = pd.read_csv("MasterVendorDevice.csv") testfile = pd.read_csv("testfile.csv", encoding='latin-1') baseDir = os.path.join("/Users", "Tim", "Desktop", "String Matcher") outDir = os.path.join(baseDir, "out") if not os.path.exists(outDir): os.makedirs(outDir) outFile = os.path.join(outDir, "matches") # 可以自定义阈值,比如这里设为80 result = StringMatch(master, testfile, score_threshold=80) print("finished")
方案二:手动过滤匹配结果(兼容旧版本FuzzyWuzzy)
如果你暂时不想升级版本,可以手动过滤每个匹配结果的分数,只保留高于阈值的项:
from fuzzywuzzy import process import pandas as pd import os def StringMatch(master, testfile, num_match=3, score_threshold=80): master_names = master.iloc[:,3] test_names = testfile.iloc[:,0] # 先获取所有匹配结果,再手动过滤分数高于阈值的 fhp_new = [] for x in test_names: matches = process.extract(x, master_names, limit=num_match) # 筛选分数>=阈值的结果 filtered_matches = [match for match in matches if match[1] >= score_threshold] # 如果过滤后结果不足num_match个,补None填充 filtered_matches += [None]*(num_match - len(filtered_matches)) fhp_new.append(filtered_matches) match_columns = [f"Match{i}" for i in range(1, num_match+1)] aggregated_matches = pd.DataFrame(fhp_new, columns=match_columns) d = {} for col in match_columns: d[col] = [y[0] if y else None for y in aggregated_matches[col]] d["test_original"] = test_names.values d["perfect match"] = d["Match1"] == d["test_original"] out = pd.DataFrame(data=d) out.to_csv(f"{outFile}.csv", index=False) return out print("starting...") master = pd.read_csv("MasterVendorDevice.csv") testfile = pd.read_csv("testfile.csv", encoding='latin-1') baseDir = os.path.join("/Users", "Tim", "Desktop", "String Matcher") outDir = os.path.join(baseDir, "out") if not os.path.exists(outDir): os.makedirs(outDir) outFile = os.path.join(outDir, "matches") result = StringMatch(master, testfile, score_threshold=80) print("finished")
关键修改点说明
- 新增了
score_threshold参数,让你可以灵活调整匹配分数的阈值 - 处理了过滤后匹配结果不足
num_match个的情况,用None填充避免报错 - 优化了列名生成逻辑,更简洁易读
- 添加了
index=False到to_csv,避免输出多余的索引列
内容的提问来源于stack exchange,提问作者isic5
相关产品推荐
相关产品推荐

