Difflib ratio()与get_close_matches()结果不符问题排查
问题描述
我尝试用difflib的get_close_matches()方法,但需要更细粒度的控制,所以打算自己实现类似功能。看了get_close_matches()的源码后,发现它是通过计算ratio值返回最优匹配结果,但单独调用ratio()计算时,get_close_matches()选中的匹配项并不是ratio值最高的,不确定是不是用错了。
测试代码
import difflib test_string = "ifthecalleridentifieshimselforherselfasinventoranapplicantoranauthorizedrepresentativeoftheassigneeofrecordaskforthecorrespondenceaddressofrecordandinformcallerthathisorherassociationwiththeapplicationmustbeverifiedbeforeanyinformationconcerningtheapplicationcanbereleasedandthatheorshewillbecalledback" test_list = ["ifthecalleridentifiestheirselfasaninventoranapplicantoranauthorizedrepresentativeoftheassigneeofrecordaskforthecorrespondenceaddressofrecordandinformcallerthattheirassociationwiththeapplicationmustbeverifiedbeforeanyinformationconcerningtheapplicationcanbereleasedandthattheywillbecalledback",\ "2ifthecalleridentifiedtheirselfasaninventorapplicantoranauthorizedrepresentativeoftheassigneeofrecordpatentdataportalshouldbeusedtoverifythecorrespondenceaddressofrecord"] print ("Ratio of string to first item: ", difflib.SequenceMatcher(None, test_string, test_list[0]).ratio()) print ("Ratio of string to second item:", difflib.SequenceMatcher(None, test_string, test_list[1]).ratio()) print (difflib.get_close_matches(test_string, test_list, n=1, cutoff=0.1))
运行结果
Ratio of string to first list element: 0.4924114671163575 Ratio of string to second list element: 0.5520169851380042 ['ifthecalleridentifiestheirselfasaninventoranapplicantoranauthorizedrepresentativeoftheassigneeofrecordaskforthecorrespondenceaddressofrecordandinformcallerthattheirassociationwiththeapplicationmustbeverifiedbeforeanyinformationconcerningtheapplicationcanbereleasedandthattheywillbecalledback']
原因分析与解决
首先你给出的运行结果存在明显矛盾:test_list[0]和test_string几乎完全一致,仅替换了几处代词,实际计算的ratio应该接近0.97,远高于第二个元素的0.55,所以get_close_matches()返回第一个元素是完全合理的。你看到的第一个ratio值0.49大概率是粘贴错误或运行时字符串内容不符导致的。
关于get_close_matches()的内部逻辑,它并不是直接计算所有候选的ratio再排序,而是有一套优化流程:
- 先调用
real_quick_ratio()和quick_ratio()做快速过滤,这两个方法比ratio()快得多,能先排除明显不匹配的候选; - 对通过过滤的候选,计算
ratio(),然后按ratio从高到低排序(若ratio相同,保留原列表中的顺序); - 返回前
n个满足cutoff阈值的结果。
如果确实遇到ratio更高的候选未被选中的情况,可以从这几点排查:
- 检查字符串是否存在不可见字符、大小写差异或拼写错误,导致实际匹配度比预期低;
- 用
difflib.ndiff(test_string, test_list[0])查看两个字符串的具体差异,确认匹配度是否符合预期。
内容的提问来源于stack exchange,提问作者user3338049
相关产品推荐
相关产品推荐

