为何fuzzywuzzy的process.extractBests未给含查询串的字符串打100分?
问题:fuzzywuzzy的process.extractBests()为何未给包含查询串的选项全打100分?
测试代码:
from fuzzywuzzy import process # Define the query string query = "Apple" # Define the list of choices choices = ["Apple", "Apple Inc.", "Apple Computer", "Apple Records", "Apple TV"] # Call the process.extractBests function results = process.extractBests(query, choices) # Print the results for result in results: print(result)
输出结果:
('Apple', 100) ('Apple Inc.', 90) ('Apple Computer', 90) ('Apple Records', 90) ('Apple TV', 90)
所有选项都包含查询串“Apple”,但评分未全为100,原因如下:
- fuzzywuzzy默认使用Levenshtein编辑距离为核心的
fuzz.WRatio作为评分器,它不是简单判断子串是否存在,而是综合考量字符串的编辑操作次数(增、删、改)和长度差异。 - 以"Apple"和"Apple Inc."为例,后者比前者多了" Inc."字符,需要通过添加操作才能匹配原查询串,编辑距离不为0,因此评分低于100;只有当两个字符串完全一致时,才会给出100分。
如果想要让所有包含查询串的选项都得到100分,可以改用fuzz.partial_ratio作为评分器,它的逻辑是判断较短字符串是否是较长字符串的子串,匹配即给100分:
修改后的代码:
from fuzzywuzzy import process, fuzz query = "Apple" choices = ["Apple", "Apple Inc.", "Apple Computer", "Apple Records", "Apple TV"] # 指定scorer为partial_ratio results = process.extractBests(query, choices, scorer=fuzz.partial_ratio) for result in results: print(result)
输出结果会变为:
('Apple', 100) ('Apple Inc.', 100) ('Apple Computer', 100) ('Apple Records', 100) ('Apple TV', 100)
内容的提问来源于stack exchange,提问作者Franck Dernoncourt
相关产品推荐
相关产品推荐

