Python如何替换字符串中命中列表的词 实现聊天敏感词过滤
聊天内容审查功能修正实现
原代码问题
原逻辑存在两个核心错误,导致无法得到预期结果:
- 匹配范围错误:判断时使用全量转小写的整句文本做匹配,只要整句包含任意违禁词,所有分词都会被替换为
***,没有做到逐词校验 - 变体适配缺失:没有处理大小写混合、字符重复拖长的变形词汇,类似
COoLlll这类写法无法被识别匹配
正确实现代码
先实现词汇归一化逻辑,统一处理大小写、连续重复的拖长字符,再逐词校验匹配:
def normalize_word(word): """统一处理词的格式:转小写 + 压缩连续重复的拖长字符""" word_lower = word.lower() if not word_lower: return "" normalized_chars = [word_lower[0]] for c in word_lower[1:]: if c != normalized_chars[-1]: normalized_chars.append(c) return "".join(normalized_chars) # 原始配置 msg = "hello bro i saw your new pc it looks really so COoLlll" banned_words = ["hello", "hi", "wow", "hmm", "cool"] # 提前归一化违禁词,用集合提升匹配效率 normalized_banned_set = {normalize_word(word) for word in banned_words} # 逐词校验替换 censored_word_list = [] for single_word in msg.split(): if normalize_word(single_word) in normalized_banned_set: censored_word_list.append("***") else: censored_word_list.append(single_word) censored_message = " ".join(censored_word_list) print(censored_message)
运行效果
执行代码后输出结果完全符合预期:
*** bro i saw your new pc it looks really so ***
该实现可以覆盖两类常见的违禁词变形场景:
- 大小写混合变体:如
HELLO、hElLo、Hi均可正常识别 - 字符拖长变体:如
hiiii、cooooooLllll、wowwwwww均可正常命中匹配
内容的提问来源于stack exchange,提问作者CB developer
相关产品推荐
相关产品推荐

