You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现保留原大小写的敏感词屏蔽?Regex是否最优?

问题:敏感词过滤时保留原文本大小写的问题

我觉得Regex是解决敏感词过滤的最佳方案,但试了两段代码都有问题:

第一段代码的问题

最初写的代码会把整个文本转成小写,不符合预期:

forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"]
def censor(string):
    # Remove line breaks and make it lowercase
    string = " ".join(string.splitlines()).lower()
    for word in forbidden_words:
        if word in string:
            string = string.replace(word, '*' * len(word))
            print(f"Forbidden word REMOVED: {word}")
    return string
print(censor("Sex, pornography, and Dicky are ALL not allowed."))

这段代码的输出:

***, ****ography, and ****y are all not allowed.

但我期望的输出是保留原文本中非敏感词的大小写:

***, ****ography, and ****y are ALL not allowed.

第二段正则代码的问题

改用正则后,又出现新问题:部分敏感词没被正确替换,比如pornography里的porn没被匹配,Dicky里的dick也没被替换:

import re

forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"]

def censor(string):
    # Remove line breaks
    string = " ".join(string.splitlines())
    for word in forbidden_words:
        # Use a regular expression to search for the word, ignoring case
        pattern = r"\b{}\b".format(word)
        if re.search(pattern, string, re.IGNORECASE):
            string = re.sub(pattern, '*' * len(word), string, flags=re.IGNORECASE)
            print(f"Forbidden word REMOVED: {word}")
    return string

print(censor("Sex, pornography, and Dicky are ALL not allowed."))

这段代码的输出:

***, pornography, and dicky are ALL not allowed.

另外想问:Regex是不是解决这个问题的最佳方案?感觉自己写的代码有点冗余,我是Python新手,求解答。


解决方案

  1. 正则匹配问题的根源
    你之前用了\b(单词边界),但porn在pornography里不是独立单词,dick在Dicky里也不是独立单词,所以\b会导致匹配失败。需要去掉\b,改用匹配子串(如果需要避免误匹配,比如drugstore里的drug是否替换,可根据需求调整规则)。

  2. 改进后的简洁代码

import re

forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"]

def censor(string):
    string = " ".join(string.splitlines())
    # 转义敏感词中的正则特殊字符,合并为单个正则模式
    pattern = re.compile('|'.join(re.escape(word) for word in forbidden_words), re.IGNORECASE)
    # 替换匹配项为对应长度的星号,同时输出被移除的敏感词
    def replace_match(match):
        matched_word = match.group()
        print(f"Forbidden word REMOVED: {matched_word.lower()}")
        return '*' * len(matched_word)
    return pattern.sub(replace_match, string)

print(censor("Sex, pornography, and Dicky are ALL not allowed."))

这段代码的输出:

Forbidden word REMOVED: sex
Forbidden word REMOVED: porn
Forbidden word REMOVED: dick
***, ****ography, and ****y are ALL not allowed.

完全符合预期。

  1. 关于Regex是否为最佳方案
    对于这种需要忽略大小写、灵活匹配子串的场景,Regex确实是高效且简洁的方案。相比循环逐个处理敏感词,合并成单个正则模式只需遍历一次文本,性能更好,代码也更简洁。之前的冗余问题就是因为循环处理每个敏感词,合并正则后即可解决。

另外需要注意:如果敏感词包含正则特殊字符(如$、.),一定要用re.escape()转义,避免正则语法错误。


内容的提问来源于stack exchange,提问作者ejade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:35:11