如何使用模糊匹配统计文本语料库中的近似字符串匹配次数
模糊匹配命中次数统计实现方案
要统计模糊匹配的命中次数,核心思路是通过滑动窗口截取长文本中与目标字符串长度接近的子串,逐一和目标做相似度比对,再对命中结果去重后统计次数,具体实现如下:
依赖安装
首先确保安装所需库:pip install fuzzywuzzy python-Levenshtein
完整代码
from fuzzywuzzy import fuzz def count_fuzzy_matches(target, text, threshold=85, overlap_tolerance=0.5): target_len = len(target) # 文本长度短于目标直接返回0 if len(text) < target_len: return 0 matches = [] # 滑动窗口逐字符截取比对 for i in range(len(text) - target_len + 1): window = text[i:i+target_len] score = fuzz.partial_token_set_ratio(target, window) if score >= threshold: matches.append((i, i+target_len)) # 去重重叠匹配 if not matches: return 0 matches.sort() unique_matches = [matches[0]] for current in matches[1:]: last = unique_matches[-1] overlap = max(0, last[1] - current[0]) # 重叠占比超过容忍度则视为同一个匹配 if overlap / (current[1] - current[0]) < overlap_tolerance: unique_matches.append(current) return len(unique_matches) # 业务测试代码 target_name = 'Example Company S.A.' text = '''This is an example text, containing various words around the company name I actually want to find. This name is Example company SA, and as you can see this will not return an exact match. Here is the name again with some modifications: example company.''' match_count = count_fuzzy_matches(target_name, text, threshold=85) print(f"模糊匹配命中次数:{match_count}") # 示例输出为2,对应两个模糊匹配的公司名
配置说明
- 阈值调整:修改
threshold参数即可调整匹配严格程度,分值范围0-100,值越高匹配要求越严格 - 重叠容忍度:
overlap_tolerance用于控制重叠匹配的去重逻辑,值为0时完全不允许重叠,值为1时允许多个匹配完全重叠,可根据业务场景调整 - 匹配算法替换:如果需要更严格的匹配规则,可以将
fuzz.partial_token_set_ratio替换为fuzz.ratio、fuzz.partial_ratio等其他fuzzywuzzy提供的算法
内容的提问来源于stack exchange,提问作者xgsktx
相关产品推荐
相关产品推荐

