如何让Discord机器人识别含间隔字符的违规词汇?
Discord机器人违规词检测优化方案
问题描述
我开发了一个Discord机器人,现有代码如下:
import discord from discord.ext import commands from replit import db intents = discord.Intents.default() intents.message_content = True intents.messages = True automod = False intents.messages = True db = {} def update_encouragements(encouraging_message): if "encouragements" in db.keys(): encouragements = db["encouragements"] encouragements.append(encouraging_message) db["encouragements"] = encouragements else: db["encouragements"] = [encouraging_message] def delete_encouragment(index): encouragements = db["encouragements"] if len(encouragements) > index: del encouragements[index] db["encouragements"] = encouragements with open("bad_wrods.txt", "r") as t: bad_words = t.read().splitlines() def get_user_mention(mention): if mention.startswith('<@') and mention.endswith('>'): return client.get_user(mention[2:-1]) @client.event async def on_message(message): await client.process_commands(message) global automod if message.author == client.user: return if message.author.bot: return if automod: if any(word in message.content.lower() for word in bad_words): await message.delete() await message.channel.send( f"Nuh uh, {message.author.mention} Please do not use bad words." ) if message.content.startswith("$automod-toggle") and message.author.id in [ 822685628201697311, 998625868643565640 ]: if automod: automod = False else: automod = True await message.channel.send( f"Automod is now {'enabled' if automod else 'disabled'}") else:
当前机器人仅能识别完整的违规词汇,我希望它能检测出字母间插入.、,、*、/、_、-及空格等字符的违规词。曾尝试在bad_words.txt中添加带空格的违规词副本,但效果不佳且不全面,请问该如何优化代码?
优化方案
核心思路是忽略干扰字符,提取消息和违规词的纯字母序列,再做匹配。具体实现如下:
修改后的完整代码
import discord from discord.ext import commands import re # 引入正则表达式库处理字符清洗 # 初始化Bot(原代码缺失此关键步骤) intents = discord.Intents.default() intents.message_content = True intents.messages = True client = commands.Bot(command_prefix="$", intents=intents) automod = False # 若不需要replit数据库功能,可注释以下代码块 # from replit import db # def update_encouragements(encouraging_message): # if "encouragements" in db.keys(): # encouragements = db["encouragements"] # encouragements.append(encouraging_message) # db["encouragements"] = encouragements # else: # db["encouragements"] = [encouraging_message] # def delete_encouragment(index): # encouragements = db["encouragements"] # if len(encouragements) > index: # del encouragements[index] # db["encouragements"] = encouragements # 加载并预处理违规词:移除干扰字符,转为纯字母小写序列 with open("bad_wrods.txt", "r") as t: bad_words = t.read().splitlines() clean_bad_words = [] for word in bad_words: # 移除所有指定干扰字符,只保留字母 clean_word = re.sub(r'[.,*/_\- ]', '', word.lower()) if clean_word: # 过滤空字符串 clean_bad_words.append(clean_word) def get_user_mention(mention): if mention.startswith('<@') and mention.endswith('>'): return client.get_user(int(mention[2:-1])) # 修复类型转换问题 @client.event async def on_message(message): await client.process_commands(message) global automod # 跳过机器人自身和其他Bot消息 if message.author == client.user or message.author.bot: return if automod: # 清洗消息内容:移除干扰字符,转为小写 cleaned_content = re.sub(r'[.,*/_\- ]', '', message.content.lower()) # 检查清洗后的内容是否包含任意违规词的纯字母序列 if any(bad_word in cleaned_content for bad_word in clean_bad_words): await message.delete() await message.channel.send(f"Nuh uh, {message.author.mention} Please do not use bad words.") # 自动审核开关命令 if message.content.startswith("$automod-toggle") and message.author.id in [822685628201697311, 998625868643565640]: automod = not automod await message.channel.send(f"Automod is now {'enabled' if automod else 'disabled'}") # 替换为你的机器人Token后启动 # client.run("YOUR_DISCORD_BOT_TOKEN")
关键优化点
- 正则字符清洗:用
re.sub()批量移除指定干扰字符,比手动替换更高效,能覆盖所有目标干扰符号。 - 双向预处理:同时对违规词和用户消息做纯字母提取,确保匹配不受干扰字符的位置和类型影响,能识别
b.a.d、b-a-d、b a d等变形违规词。 - 修复原代码问题:补上了Bot初始化步骤、修复了用户ID转换的类型错误、删除了重复的
intents设置,修正了语法错误。
内容的提问来源于stack exchange,提问作者John Hawking
相关产品推荐
相关产品推荐

