You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过相似度匹配判断特定字符串是否包含在句子中

实现方案

环境准备

首先安装所需依赖库:

  • 核心匹配库用你提到的jellyfish,Windows平台读取本地Outlook邮件使用pywin32
  • 安装命令:
    pip install jellyfish pywin32

注:如果不需要读取本地Outlook客户端邮件,仅处理已有的主题文本,可以不用安装pywin32;非Windows平台可以用imaplib直接连接邮件服务器拉取主题。

核心相似度匹配逻辑实现

你示例中用到的是Jaro-Winkler相似度算法,jellyfish原生支持该算法,刚好符合需求,匹配逻辑为:将目标句子拆分为单个单词,遍历每个单词和预设关键词计算相似度,取最高值和阈值对比判断是否匹配。

import jellyfish

def get_max_similarity(target_sentence: str, keyword: str) -> float:
    # 统一转小写,移除标点符号避免干扰
    words = [word.strip('.,!?()[]{}"\'').lower() for word in target_sentence.split()]
    keyword_lower = keyword.lower()
    max_sim = 0.0
    for word in words:
        sim = jellyfish.jaro_winkler_similarity(word, keyword_lower)
        if sim > max_sim:
            max_sim = sim
    return max_sim

你给出的示例测试可以直接复用该函数:
get_max_similarity("The community is here to help you with specific coding, algorithm, or language problems.", "algorism")
输出结果为0.8248242811501597,和你示例的输出完全一致。

分类规则配置

你可以根据业务需求自定义分类、对应关键词列表和相似度阈值,阈值建议取值范围为0.7-0.8,可根据实际匹配效果调整:

# 配置规则结构:分类名称: [关键词列表, 相似度阈值]
CLASSIFY_RULES = {
    "技术需求": [["coding", "algorithm", "bug", "debug"], 0.75],
    "财务相关": [["invoice", "payment", "bill", "报销"], 0.7],
    "人事通知": [["offer", "入职", "离职", "绩效"], 0.7]
}

Outlook邮件主题读取与分类实现

import win32com.client

def classify_outlook_subjects(read_count: int = 100) -> list:
    # 连接本地Outlook客户端
    outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI")
    # 3为收件箱的默认索引
    inbox = outlook.GetDefaultFolder(3)
    # 按收件时间倒序排序,优先读取最新邮件
    mails = inbox.Items
    mails.Sort("[ReceivedTime]", True)

    result = []
    for idx, mail in enumerate(mails):
        if idx >= read_count:
            break
        subject = mail.Subject
        final_class = "未分类"
        match_keyword = None
        match_sim = 0.0
        # 遍历所有分类规则匹配
        for class_name, (keyword_list, threshold) in CLASSIFY_RULES.items():
            for keyword in keyword_list:
                current_sim = get_max_similarity(subject, keyword)
                if current_sim >= threshold and current_sim > match_sim:
                    match_sim = current_sim
                    final_class = class_name
                    match_keyword = keyword
        result.append({
            "邮件主题": subject,
            "分类结果": final_class,
            "匹配关键词": match_keyword,
            "相似度": round(match_sim, 4) if match_sim > 0 else None
        })
    return result

运行测试

if __name__ == "__main__":
    # 读取最新50封邮件分类
    res = classify_outlook_subjects(read_count=50)
    for item in res:
        print(item)

可选调整项

  • 如果处理中文主题,建议先用jieba分词库对中文句子拆分后再计算相似度,Jaro-Winkler算法对中英文字符均兼容
  • 对准确率要求高的场景可以把相似度阈值上调到0.8以上,需要降低漏判率可以下调到0.7左右
  • 也可以根据需求替换为jellyfish的其他相似度算法,比如编辑距离、汉明距离等

内容的提问来源于stack exchange,提问作者조현성

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 14:54:10