如何在Python中实现保留原大小写的敏感词屏蔽?Regex是否最优?
问题:敏感词过滤时保留原文本大小写的问题
我觉得Regex是解决敏感词过滤的最佳方案,但试了两段代码都有问题:
第一段代码的问题
最初写的代码会把整个文本转成小写,不符合预期:
forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"] def censor(string): # Remove line breaks and make it lowercase string = " ".join(string.splitlines()).lower() for word in forbidden_words: if word in string: string = string.replace(word, '*' * len(word)) print(f"Forbidden word REMOVED: {word}") return string print(censor("Sex, pornography, and Dicky are ALL not allowed."))
这段代码的输出:
***, ****ography, and ****y are all not allowed.
但我期望的输出是保留原文本中非敏感词的大小写:
***, ****ography, and ****y are ALL not allowed.
第二段正则代码的问题
改用正则后,又出现新问题:部分敏感词没被正确替换,比如pornography里的porn没被匹配,Dicky里的dick也没被替换:
import re forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"] def censor(string): # Remove line breaks string = " ".join(string.splitlines()) for word in forbidden_words: # Use a regular expression to search for the word, ignoring case pattern = r"\b{}\b".format(word) if re.search(pattern, string, re.IGNORECASE): string = re.sub(pattern, '*' * len(word), string, flags=re.IGNORECASE) print(f"Forbidden word REMOVED: {word}") return string print(censor("Sex, pornography, and Dicky are ALL not allowed."))
这段代码的输出:
***, pornography, and dicky are ALL not allowed.
另外想问:Regex是不是解决这个问题的最佳方案?感觉自己写的代码有点冗余,我是Python新手,求解答。
解决方案
正则匹配问题的根源
你之前用了\b(单词边界),但porn在pornography里不是独立单词,dick在Dicky里也不是独立单词,所以\b会导致匹配失败。需要去掉\b,改用匹配子串(如果需要避免误匹配,比如drugstore里的drug是否替换,可根据需求调整规则)。改进后的简洁代码
import re forbidden_words = ["sex", "porn", "dick", "drug", "casino", "gambling"] def censor(string): string = " ".join(string.splitlines()) # 转义敏感词中的正则特殊字符,合并为单个正则模式 pattern = re.compile('|'.join(re.escape(word) for word in forbidden_words), re.IGNORECASE) # 替换匹配项为对应长度的星号,同时输出被移除的敏感词 def replace_match(match): matched_word = match.group() print(f"Forbidden word REMOVED: {matched_word.lower()}") return '*' * len(matched_word) return pattern.sub(replace_match, string) print(censor("Sex, pornography, and Dicky are ALL not allowed."))
这段代码的输出:
Forbidden word REMOVED: sex Forbidden word REMOVED: porn Forbidden word REMOVED: dick ***, ****ography, and ****y are ALL not allowed.
完全符合预期。
- 关于Regex是否为最佳方案
对于这种需要忽略大小写、灵活匹配子串的场景,Regex确实是高效且简洁的方案。相比循环逐个处理敏感词,合并成单个正则模式只需遍历一次文本,性能更好,代码也更简洁。之前的冗余问题就是因为循环处理每个敏感词,合并正则后即可解决。
另外需要注意:如果敏感词包含正则特殊字符(如$、.),一定要用re.escape()转义,避免正则语法错误。
内容的提问来源于stack exchange,提问作者ejade
相关产品推荐
相关产品推荐

