如何实现忽略指定词语及特殊字符的子字符串匹配功能
解决方案
实现思路
- 核心逻辑为「先清洗文本过滤无效内容,再做有序子序列匹配」,比直接编写复杂正则可维护性更高,20个指定忽略词也可以快速配置调整
分步实现方案
步骤1:预定义配置
- 提取待匹配短语的有序关键词列表,比如
"cat fish"处理为phrase_keywords = ["cat", "fish"] - 把需要跳过的指定词语存入集合,方便O(1)效率快速判断:
stop_words = {"likes", ...} /* 剩余19个忽略词自行补充 */
步骤2:文本清洗处理
对待检测文本做两层过滤:
- 过滤所有特殊字符、标点、HTML标签,仅保留有效单词字符(英文/汉字)和空格
- 拆分文本为单词列表,删除所有属于
stop_words的词语,得到纯有效词序列
步骤3:匹配判断
检查phrase_keywords是否是纯有效词序列的连续子序列,是则返回True,否则返回False
正则快速实现方案(适合简单场景)
如果不想写完整清洗逻辑,可直接生成适配的正则表达式:
两个关键词之间的匹配规则写为 (?:[\W_]|{{停用词}})*,其中[\W_]匹配所有非单词字符、标点、特殊符号,*允许中间内容出现0次到多次。
对应示例的正则为:
\bcat(?:[\W_]|likes)*fish\b
开启忽略大小写标志后可适配大小写不敏感的匹配场景,你给出的前4个示例均可匹配成功,第5个示例因中间存在未被忽略的and、eats词语无法匹配,完全符合要求。
Python完整实现代码
import re def match_phrase(text: str, phrase: str, stop_words: set) -> bool: # 1. 预处理待匹配短语 phrase_keywords = phrase.lower().split() # 2. 清洗文本:去除所有非单词字符,拆分后过滤停用词 cleaned_text = re.sub(r'[^\w\s]', '', text.lower()) text_words = [w for w in cleaned_text.split() if w not in stop_words] # 3. 匹配连续子序列 n = len(phrase_keywords) for i in range(len(text_words) - n + 1): if text_words[i:i+n] == phrase_keywords: return True return False # 测试用例 stop_words = {"likes"} phrase = "cat fish" test_cases = [ "There is a <strong>cat fish</strong>.", "There is a <strong>cat, fish</strong> and a dog.", "My <strong>cat <em>likes</em> fish</strong> very much.", "My <strong>cat <em>likes</em>)- fish</strong> very much.", "My cat likes and eats fish a lot." ] for case in test_cases: print(match_phrase(case, phrase, stop_words)) # 输出:True True True True False,完全符合示例要求
内容的提问来源于stack exchange,提问作者Miroslav Georgiev
相关产品推荐
相关产品推荐

