You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python正则表达式敏感词检测中为屏蔽词添加例外规则

该需求完全可以通过正则实现,核心是利用负向零宽断言排除豁免场景,或者在替换回调中增加规则判断,以下是具体实现方案:

方案一:纯正则实现(适合规则简单固定的场景)

通过在正则中添加负向零宽断言,直接把两类豁免场景排除在匹配结果之外,不需要修改替换逻辑:

  1. 添加(?!(?-i:BAR)\b)断言:仅匹配全大写的BAR完整单词时触发,跳过该内容的匹配,(?-i:)是内联修饰符,保证该断言不受全局re.IGNORECASE的大小写忽略影响
  2. 添加(?!foo(?=\s+good\b))断言:匹配到foo后紧跟空格加good完整短语时触发,跳过该内容的匹配

完整实现代码如下:

import re

profanity_list = ['foo', 'bar']
pattern_profanity = re.compile(
    r'\b(?!(?-i:BAR)\b)(?!foo(?=\s+good\b))({})\b'.format('|'.join(profanity_list)),
    flags=re.IGNORECASE
)
s = 'foo BAR foo good Bar'
censor_char = '*'
result = pattern_profanity.sub(repl=lambda m: censor_char*len(m.group(0)), string=s)
print(result)
# 输出结果:*** BAR foo good ***

方案二:替换回调加规则判断(适合规则多、后续需要频繁调整的场景)

如果后续要新增更多豁免规则,全写在正则里会导致可读性变差,可以先匹配所有敏感词,在替换回调函数中判断是否属于豁免场景,属于则返回原内容,否则返回屏蔽字符:

import re

profanity_list = ['foo', 'bar']
pattern_profanity = re.compile(
    r'\b({})\b'.format('|'.join(profanity_list)),
    flags=re.IGNORECASE
)

def censor_replace(match):
    matched_str = match.group(0)
    # 豁免规则1:全大写的BAR
    if matched_str == 'BAR':
        return matched_str
    # 豁免规则2:foo 后面紧跟 good 的情况
    if matched_str.lower() == 'foo':
        end_pos = match.end()
        # 校验匹配结束位置后是否为空格加good
        if len(s) > end_pos and s[end_pos:].lstrip().startswith('good'):
            return matched_str
    # 非豁免场景返回屏蔽字符
    return '*' * len(matched_str)

s = 'foo BAR foo good Bar'
result = pattern_profanity.sub(repl=censor_replace, string=s)
print(result)
# 输出结果:*** BAR foo good ***

内容的提问来源于stack exchange,提问作者CX_Lin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 04:30:01