You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用模糊匹配统计文本语料库中的近似字符串匹配次数

模糊匹配命中次数统计实现方案

要统计模糊匹配的命中次数,核心思路是通过滑动窗口截取长文本中与目标字符串长度接近的子串,逐一和目标做相似度比对,再对命中结果去重后统计次数,具体实现如下:

依赖安装

首先确保安装所需库:
pip install fuzzywuzzy python-Levenshtein

完整代码

from fuzzywuzzy import fuzz

def count_fuzzy_matches(target, text, threshold=85, overlap_tolerance=0.5):
    target_len = len(target)
    # 文本长度短于目标直接返回0
    if len(text) < target_len:
        return 0
    matches = []
    # 滑动窗口逐字符截取比对
    for i in range(len(text) - target_len + 1):
        window = text[i:i+target_len]
        score = fuzz.partial_token_set_ratio(target, window)
        if score >= threshold:
            matches.append((i, i+target_len))
    # 去重重叠匹配
    if not matches:
        return 0
    matches.sort()
    unique_matches = [matches[0]]
    for current in matches[1:]:
        last = unique_matches[-1]
        overlap = max(0, last[1] - current[0])
        # 重叠占比超过容忍度则视为同一个匹配
        if overlap / (current[1] - current[0]) < overlap_tolerance:
            unique_matches.append(current)
    return len(unique_matches)

# 业务测试代码
target_name = 'Example Company S.A.'
text = '''This is an example text, containing various
words around the company
name I actually want to find. This name is Example
company SA, and as you can
see this will not return an exact match. Here is the
name again with some modifications: example company.'''

match_count = count_fuzzy_matches(target_name, text, threshold=85)
print(f"模糊匹配命中次数:{match_count}") # 示例输出为2,对应两个模糊匹配的公司名

配置说明

  • 阈值调整:修改threshold参数即可调整匹配严格程度,分值范围0-100,值越高匹配要求越严格
  • 重叠容忍度:overlap_tolerance用于控制重叠匹配的去重逻辑,值为0时完全不允许重叠,值为1时允许多个匹配完全重叠,可根据业务场景调整
  • 匹配算法替换:如果需要更严格的匹配规则,可以将fuzz.partial_token_set_ratio替换为fuzz.ratio、fuzz.partial_ratio等其他fuzzywuzzy提供的算法

内容的提问来源于stack exchange,提问作者xgsktx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 04:36:08