You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现忽略指定词语及特殊字符的子字符串匹配功能

解决方案

实现思路

  • 核心逻辑为「先清洗文本过滤无效内容,再做有序子序列匹配」,比直接编写复杂正则可维护性更高,20个指定忽略词也可以快速配置调整

分步实现方案

步骤1:预定义配置

  • 提取待匹配短语的有序关键词列表,比如"cat fish"处理为 phrase_keywords = ["cat", "fish"]
  • 把需要跳过的指定词语存入集合,方便O(1)效率快速判断:stop_words = {"likes", ...} /* 剩余19个忽略词自行补充 */

步骤2:文本清洗处理

对待检测文本做两层过滤:

  1. 过滤所有特殊字符、标点、HTML标签,仅保留有效单词字符(英文/汉字)和空格
  2. 拆分文本为单词列表,删除所有属于stop_words的词语,得到纯有效词序列

步骤3:匹配判断

检查phrase_keywords是否是纯有效词序列的连续子序列,是则返回True,否则返回False

正则快速实现方案(适合简单场景)

如果不想写完整清洗逻辑,可直接生成适配的正则表达式:
两个关键词之间的匹配规则写为 (?:[\W_]|{{停用词}})*,其中[\W_]匹配所有非单词字符、标点、特殊符号,*允许中间内容出现0次到多次。
对应示例的正则为:

\bcat(?:[\W_]|likes)*fish\b

开启忽略大小写标志后可适配大小写不敏感的匹配场景,你给出的前4个示例均可匹配成功,第5个示例因中间存在未被忽略的and、eats词语无法匹配,完全符合要求。

Python完整实现代码

import re

def match_phrase(text: str, phrase: str, stop_words: set) -> bool:
    # 1. 预处理待匹配短语
    phrase_keywords = phrase.lower().split()
    # 2. 清洗文本:去除所有非单词字符,拆分后过滤停用词
    cleaned_text = re.sub(r'[^\w\s]', '', text.lower())
    text_words = [w for w in cleaned_text.split() if w not in stop_words]
    # 3. 匹配连续子序列
    n = len(phrase_keywords)
    for i in range(len(text_words) - n + 1):
        if text_words[i:i+n] == phrase_keywords:
            return True
    return False

# 测试用例
stop_words = {"likes"}
phrase = "cat fish"
test_cases = [
    "There is a <strong>cat fish</strong>.",
    "There is a <strong>cat, fish</strong> and a dog.",
    "My <strong>cat <em>likes</em> fish</strong> very much.",
    "My <strong>cat <em>likes</em>)- fish</strong> very much.",
    "My cat likes and eats fish a lot."
]

for case in test_cases:
    print(match_phrase(case, phrase, stop_words))
# 输出:True True True True False,完全符合示例要求

内容的提问来源于stack exchange,提问作者Miroslav Georgiev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 00:12:03