You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Fuzzywuzzy字符串匹配项目中添加分数阈值过滤低匹配结果?

解决FuzzyWuzzy中score_cutoff参数不生效及添加匹配分数阈值的问题

我来帮你搞定这个匹配分数阈值的需求!你遇到process.extract不接受score_cutoff参数的问题,多半是因为使用的fuzzywuzzy版本比较旧——这个参数是在后续版本中新增的特性。下面我给你两种解决方案,不管是升级版本还是手动过滤,都能实现只保留高分匹配结果的目标。

方案一:升级FuzzyWuzzy并使用score_cutoff参数

首先先把fuzzywuzzy升级到最新版本,建议同时安装python-Levenshtein来大幅提升匹配速度:

pip install --upgrade fuzzywuzzy python-Levenshtein

接下来修改你的代码,添加score_cutoff参数(比如设为80,你可以根据需求调整),同时修正原代码里的一些小问题(比如limit参数和函数的num_match不统一、列名生成逻辑优化等):

from fuzzywuzzy import process
import pandas as pd
import os

def StringMatch(master, testfile, num_match=3, score_threshold=80):
    master_names = master.iloc[:,3]
    test_names = testfile.iloc[:,0]
    
    # 加入score_cutoff参数,只保留分数高于阈值的结果
    fhp_new = [process.extract(x, master_names, limit=num_match, score_cutoff=score_threshold) for x in test_names]
    
    # 生成列名,更简洁的方式
    match_columns = [f"Match{i}" for i in range(1, num_match+1)]
    aggregated_matches = pd.DataFrame(fhp_new, columns=match_columns)
    
    d = {}
    for col in match_columns:
        # 处理可能的空匹配(如果没有达到阈值的结果,填充为None)
        d[col] = [y[0] if y else None for y in aggregated_matches[col]]
    
    d["test_original"] = test_names.values
    # 完美匹配判断也要考虑空值的情况
    d["perfect match"] = d["Match1"] == d["test_original"]
    
    out = pd.DataFrame(data=d)
    out.to_csv(f"{outFile}.csv", index=False)
    return out

print("starting...")
master = pd.read_csv("MasterVendorDevice.csv")
testfile = pd.read_csv("testfile.csv", encoding='latin-1')
baseDir = os.path.join("/Users", "Tim", "Desktop", "String Matcher")
outDir = os.path.join(baseDir, "out")
if not os.path.exists(outDir):
    os.makedirs(outDir)
outFile = os.path.join(outDir, "matches")
# 可以自定义阈值,比如这里设为80
result = StringMatch(master, testfile, score_threshold=80)
print("finished")

方案二:手动过滤匹配结果(兼容旧版本FuzzyWuzzy)

如果你暂时不想升级版本,可以手动过滤每个匹配结果的分数,只保留高于阈值的项:

from fuzzywuzzy import process
import pandas as pd
import os

def StringMatch(master, testfile, num_match=3, score_threshold=80):
    master_names = master.iloc[:,3]
    test_names = testfile.iloc[:,0]
    
    # 先获取所有匹配结果,再手动过滤分数高于阈值的
    fhp_new = []
    for x in test_names:
        matches = process.extract(x, master_names, limit=num_match)
        # 筛选分数>=阈值的结果
        filtered_matches = [match for match in matches if match[1] >= score_threshold]
        # 如果过滤后结果不足num_match个,补None填充
        filtered_matches += [None]*(num_match - len(filtered_matches))
        fhp_new.append(filtered_matches)
    
    match_columns = [f"Match{i}" for i in range(1, num_match+1)]
    aggregated_matches = pd.DataFrame(fhp_new, columns=match_columns)
    
    d = {}
    for col in match_columns:
        d[col] = [y[0] if y else None for y in aggregated_matches[col]]
    
    d["test_original"] = test_names.values
    d["perfect match"] = d["Match1"] == d["test_original"]
    
    out = pd.DataFrame(data=d)
    out.to_csv(f"{outFile}.csv", index=False)
    return out

print("starting...")
master = pd.read_csv("MasterVendorDevice.csv")
testfile = pd.read_csv("testfile.csv", encoding='latin-1')
baseDir = os.path.join("/Users", "Tim", "Desktop", "String Matcher")
outDir = os.path.join(baseDir, "out")
if not os.path.exists(outDir):
    os.makedirs(outDir)
outFile = os.path.join(outDir, "matches")
result = StringMatch(master, testfile, score_threshold=80)
print("finished")

关键修改点说明

  • 新增了score_threshold参数,让你可以灵活调整匹配分数的阈值
  • 处理了过滤后匹配结果不足num_match个的情况,用None填充避免报错
  • 优化了列名生成逻辑,更简洁易读
  • 添加了index=False到to_csv,避免输出多余的索引列

内容的提问来源于stack exchange,提问作者isic5

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:35:33