如何优化FuzzyWuzzy匹配逻辑,精准识别用户支持请求的关联应用?
问题背景
公司用户每日在工作频道提交多份支持请求,涉及Okta、Slack、Cloudflare等应用,希望通过Bot自动推送对应归档的解决方案。当前采用Python的FuzzyWuzzy库,用fuzz.partial_token_set_ratio计算用户消息与预设请求类型(如OKTA_ACCESS、OKTA_UNLOCK、CLOUDFLARE)的匹配分数,设定分数>90时判定匹配。
但现有逻辑存在误匹配问题:当用户发送如下消息时,Cloudflare和Okta的匹配分数均>90,但实际用户问题仅与Cloudflare相关:
Hello team, it looks like I am having an issue with Cloudflare. I have been added to the following Okta group.
Can someone look into this for me, please?
现有代码如下:
from enum import Enum from fuzzywuzzy import fuzz class RequestType(Enum): OKTA_ACCESS = "okta access" OKTA_UNLOCK = "okta unlock" CLOUDFLARE = "cannot connect to cloudflare" class RequestsMatcher: def __init__(self, message: str) -> None: scores = [ ( type, fuzz.partial_token_set_ratio(message, type.value) ) for type in RequestType ] mvs = [score for score in scores if score[1] > 75] hits = [score for score in scores if score[1] > 90] print(scores) self.scores = scores self.highest = mvs self.hit = next(iter(hits), None)
优化方案
1. 聚焦问题分句匹配
用户核心诉求通常和"issue"、"problem"、"cannot"这类问题描述词绑定,先提取包含这类词汇的分句,仅针对分句做匹配,避免无关内容干扰。
修改后示例代码:
import re from enum import Enum from fuzzywuzzy import fuzz class RequestType(Enum): OKTA_ACCESS = "okta access" OKTA_UNLOCK = "okta unlock" CLOUDFLARE = "cloudflare" # 简化匹配文本,聚焦应用核心标识 class RequestsMatcher: def __init__(self, message: str) -> None: # 提取包含问题关键词的分句 issue_keywords = r"(?i).*(issue|problem|cannot|trouble|unable|issue with).*" issue_sentences = [sent.strip() for sent in message.split('.') if re.match(issue_keywords, sent)] # 优先用问题分句匹配,无匹配时 fallback 到全消息 target_text = ' '.join(issue_sentences).lower() if issue_sentences else message.lower() scores = [ (req_type, fuzz.partial_token_set_ratio(target_text, req_type.value.lower())) for req_type in RequestType ] # 按分数排序,取最高分判断是否达标 sorted_scores = sorted(scores, key=lambda x: x[1], reverse=True) top_score = sorted_scores[0][1] self.hit = sorted_scores[0] if top_score > 90 else None print(scores) self.scores = scores
2. 简化匹配模板+分层阈值
把RequestType的value简化为"应用名+核心动作"的极简形式,比如CLOUDFLARE = "cloudflare"、OKTA_UNLOCK = "okta unlock",减少冗余文本带来的误匹配。同时给不同类型设置分层阈值:
- 仅应用名的匹配(如Cloudflare):阈值设为90
- 带动作的匹配(如okta unlock):阈值设为95,避免单纯提到应用名就触发匹配
3. 关键词位置加权
计算分数时,给靠近问题词汇的应用名额外加权。比如检测到应用名出现在"issue"、"problem"前后3个词范围内,分数加15分;若出现在无关陈述中(如示例里的"I have been added to the following Okta group"),分数减10分。
4. 切换至意图分类方案
如果模糊匹配精度始终达不到要求,可基于历史支持请求数据,用spaCy做实体识别+意图分类,或用scikit-learn训练朴素贝叶斯/逻辑回归分类模型,通过标注数据让模型学习用户诉求与应用的关联,能大幅提升匹配准确率。
内容的提问来源于stack exchange,提问作者Mervin Hemaraju

