如何通过相似度匹配判断特定字符串是否包含在句子中
实现方案
环境准备
首先安装所需依赖库:
- 核心匹配库用你提到的
jellyfish,Windows平台读取本地Outlook邮件使用pywin32 - 安装命令:
pip install jellyfish pywin32
注:如果不需要读取本地Outlook客户端邮件,仅处理已有的主题文本,可以不用安装
pywin32;非Windows平台可以用imaplib直接连接邮件服务器拉取主题。
核心相似度匹配逻辑实现
你示例中用到的是Jaro-Winkler相似度算法,jellyfish原生支持该算法,刚好符合需求,匹配逻辑为:将目标句子拆分为单个单词,遍历每个单词和预设关键词计算相似度,取最高值和阈值对比判断是否匹配。
import jellyfish def get_max_similarity(target_sentence: str, keyword: str) -> float: # 统一转小写,移除标点符号避免干扰 words = [word.strip('.,!?()[]{}"\'').lower() for word in target_sentence.split()] keyword_lower = keyword.lower() max_sim = 0.0 for word in words: sim = jellyfish.jaro_winkler_similarity(word, keyword_lower) if sim > max_sim: max_sim = sim return max_sim
你给出的示例测试可以直接复用该函数:get_max_similarity("The community is here to help you with specific coding, algorithm, or language problems.", "algorism")
输出结果为0.8248242811501597,和你示例的输出完全一致。
分类规则配置
你可以根据业务需求自定义分类、对应关键词列表和相似度阈值,阈值建议取值范围为0.7-0.8,可根据实际匹配效果调整:
# 配置规则结构:分类名称: [关键词列表, 相似度阈值] CLASSIFY_RULES = { "技术需求": [["coding", "algorithm", "bug", "debug"], 0.75], "财务相关": [["invoice", "payment", "bill", "报销"], 0.7], "人事通知": [["offer", "入职", "离职", "绩效"], 0.7] }
Outlook邮件主题读取与分类实现
import win32com.client def classify_outlook_subjects(read_count: int = 100) -> list: # 连接本地Outlook客户端 outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI") # 3为收件箱的默认索引 inbox = outlook.GetDefaultFolder(3) # 按收件时间倒序排序,优先读取最新邮件 mails = inbox.Items mails.Sort("[ReceivedTime]", True) result = [] for idx, mail in enumerate(mails): if idx >= read_count: break subject = mail.Subject final_class = "未分类" match_keyword = None match_sim = 0.0 # 遍历所有分类规则匹配 for class_name, (keyword_list, threshold) in CLASSIFY_RULES.items(): for keyword in keyword_list: current_sim = get_max_similarity(subject, keyword) if current_sim >= threshold and current_sim > match_sim: match_sim = current_sim final_class = class_name match_keyword = keyword result.append({ "邮件主题": subject, "分类结果": final_class, "匹配关键词": match_keyword, "相似度": round(match_sim, 4) if match_sim > 0 else None }) return result
运行测试
if __name__ == "__main__": # 读取最新50封邮件分类 res = classify_outlook_subjects(read_count=50) for item in res: print(item)
可选调整项
- 如果处理中文主题,建议先用
jieba分词库对中文句子拆分后再计算相似度,Jaro-Winkler算法对中英文字符均兼容 - 对准确率要求高的场景可以把相似度阈值上调到0.8以上,需要降低漏判率可以下调到0.7左右
- 也可以根据需求替换为jellyfish的其他相似度算法,比如编辑距离、汉明距离等
内容的提问来源于stack exchange,提问作者조현성
相关产品推荐
相关产品推荐

