如何修改Python代码实现大小写不敏感的敏感词替换(禁用for/in)
敏感词替换问题修正方案
问题描述
现有敏感词列表 banned=["things", "show","nature","strange image"],需实现:
- 读取论坛消息文本文件,将所有敏感词替换为等长星号
- 敏感词匹配不区分大小写(如
Show需替换,shown不应被替换) - 禁止使用
for和in语句
当前代码存在两个问题:
- 无法替换
Show这类大小写变体的敏感词 - 会误替换
shown中包含的show片段
原代码如下:
#banned_list banned=["things","nature","strange image","show"] #read message with open("forum1","r")as f: message=f.readlines() #append modified message in a new list new_forum=[] i=0 while i<len(message): j=0 while j<len(banned): if message[i].__contains__(banned[j]): message[i]=message[i].replace(banned[j],len(banned[j])*"*") j+=1 else: j+=1 new_forum.append(message[i]) i+=1 #write to a new_list with open("new_forum1","w")as n: i=0 while i<len(new_forum): n.write(new_forum[i]) i+=1
修改思路
- 解决大小写匹配问题:使用正则表达式的不区分大小写模式,替代原生字符串的精确匹配
- 避免片段误替换:用正则边界规则确保只匹配完整敏感词,不会命中包含敏感词片段的单词
- 规避for/in语句:全程用while循环遍历文本行和敏感词列表,符合作业限制
修改后的代码
import re # 敏感词列表 banned = ["things", "nature", "strange image", "show"] # 读取消息文件 with open("forum1", "r") as f: message = f.readlines() new_forum = [] i = 0 while i < len(message): current_line = message[i] j = 0 while j < len(banned): word = banned[j] # 构造正则规则:转义敏感词特殊字符,匹配非单词字符边界,不区分大小写 pattern = re.compile(r'(?<!\w)' + re.escape(word) + r'(?!\w)', re.IGNORECASE) # 替换为等长星号 current_line = pattern.sub(len(word)*'*', current_line) j += 1 new_forum.append(current_line) i += 1 # 写入新文件 with open("new_forum1", "w") as n: i = 0 while i < len(new_forum): n.write(new_forum[i]) i += 1
代码说明
re.escape(word):转义敏感词中的特殊字符,避免正则语法错误(?<!\w)和(?!\w):确保敏感词前后是非单词字符,仅匹配完整独立的敏感词,避免误替换shown这类单词re.IGNORECASE:开启不区分大小写匹配,覆盖Show、SHOW等所有大小写变体- 全程使用while循环遍历,未使用
for和in语句,符合作业要求
内容的提问来源于stack exchange,提问作者Lu Yubing
相关产品推荐
相关产品推荐

