如何在嵌套循环中为外层循环每个元素获取最大匹配值?
纯Python实现模糊匹配最大匹配度筛选
需求背景
现有两个词语列表,使用fuzzywuzzy库的fuzz.ratio方法计算词语间的匹配度,当前通过嵌套循环得到所有组合的匹配结果:
from fuzzywuzzy import fuzz l = ['mango','apple'] l2 = ['ola','john'] for i in l: for j in l2: print(i,j,fuzz.ratio(i,j))
输出结果:
mango ola 25 mango john 22 apple ola 25 apple john 0
需要为外层列表的每个元素,筛选出与之匹配度最高的结果,预期输出:
mango ola 25 apple ola 25
已知可通过pandas实现该需求,以下提供纯Python的实现方案,同时附上参考的pandas实现代码。
纯Python实现方案
核心思路是针对外层列表的每个元素,遍历内层列表计算匹配度,实时记录当前元素的最大匹配度及对应内层元素,最后汇总结果输出:
from fuzzywuzzy import fuzz l = ['mango','apple'] l2 = ['ola','john'] max_match_results = [] for outer_word in l: # 初始化最大匹配度为-1(匹配度范围为0-100,确保第一个结果能被捕获) highest_ratio = -1 matched_word = None for inner_word in l2: current_ratio = fuzz.ratio(outer_word, inner_word) # 若当前匹配度更高,则更新记录 if current_ratio > highest_ratio: highest_ratio = current_ratio matched_word = inner_word # 若需保留所有匹配度等于最大值的结果,可改为列表存储并判断相等的情况 max_match_results.append((outer_word, matched_word, highest_ratio)) # 打印最终结果 for res in max_match_results: print(res[0], res[1], res[2])
运行后即可得到预期的输出结果。
参考的pandas实现方案
通过将所有匹配结果存入DataFrame,分组后筛选每组中匹配度最大的行:
from fuzzywuzzy import fuzz import pandas as pd l = ['mango','apple'] l2 = ['ola','john'] result = [] for i in l: for j in l2: result.append((i, j, fuzz.ratio(i, j))) # 转换为DataFrame并指定列名 df = pd.DataFrame(result, columns=['word1', 'word2', 'distance']) # 分组后标记出每组中匹配度等于最大值的行 idx = df.groupby(['word1'])['distance'].transform(max) == df['distance'] print(df[idx])
内容的提问来源于stack exchange,提问作者moth
相关产品推荐
相关产品推荐

