You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Discord机器人识别含间隔字符的违规词汇?

Discord机器人违规词检测优化方案

问题描述

我开发了一个Discord机器人,现有代码如下:

import discord
from discord.ext import commands
from replit import db

intents = discord.Intents.default()
intents.message_content = True
intents.messages = True

automod = False
intents.messages = True

db = {}
def update_encouragements(encouraging_message):
    if "encouragements" in db.keys():
        encouragements = db["encouragements"]
        encouragements.append(encouraging_message)
        db["encouragements"] = encouragements
    else:
        db["encouragements"] = [encouraging_message]

def delete_encouragment(index):
    encouragements = db["encouragements"]
    if len(encouragements) > index:
        del encouragements[index]
    db["encouragements"] = encouragements

with open("bad_wrods.txt", "r") as t:
    bad_words = t.read().splitlines()

def get_user_mention(mention):
    if mention.startswith('<@') and mention.endswith('>'):
        return client.get_user(mention[2:-1])

@client.event
async def on_message(message):
    await client.process_commands(message)
    global automod
    if message.author == client.user:
        return

    if message.author.bot:
        return
    if automod:
        if any(word in message.content.lower() for word in bad_words):
            await message.delete()
            await message.channel.send(
                f"Nuh uh, {message.author.mention} Please do not use bad words."
            )

    if message.content.startswith("$automod-toggle") and message.author.id in [
            822685628201697311, 998625868643565640
    ]:
        if automod:
            automod = False
        else:
            automod = True
        await message.channel.send(
            f"Automod is now {'enabled' if automod else 'disabled'}")
        
        else:

当前机器人仅能识别完整的违规词汇,我希望它能检测出字母间插入.、,、*、/、_、-及空格等字符的违规词。曾尝试在bad_words.txt中添加带空格的违规词副本,但效果不佳且不全面,请问该如何优化代码?

优化方案

核心思路是忽略干扰字符,提取消息和违规词的纯字母序列,再做匹配。具体实现如下:

修改后的完整代码

import discord
from discord.ext import commands
import re  # 引入正则表达式库处理字符清洗

# 初始化Bot(原代码缺失此关键步骤)
intents = discord.Intents.default()
intents.message_content = True
intents.messages = True
client = commands.Bot(command_prefix="$", intents=intents)

automod = False

# 若不需要replit数据库功能,可注释以下代码块
# from replit import db
# def update_encouragements(encouraging_message):
#     if "encouragements" in db.keys():
#         encouragements = db["encouragements"]
#         encouragements.append(encouraging_message)
#         db["encouragements"] = encouragements
#     else:
#         db["encouragements"] = [encouraging_message]

# def delete_encouragment(index):
#     encouragements = db["encouragements"]
#     if len(encouragements) > index:
#         del encouragements[index]
#     db["encouragements"] = encouragements

# 加载并预处理违规词:移除干扰字符,转为纯字母小写序列
with open("bad_wrods.txt", "r") as t:
    bad_words = t.read().splitlines()
clean_bad_words = []
for word in bad_words:
    # 移除所有指定干扰字符,只保留字母
    clean_word = re.sub(r'[.,*/_\- ]', '', word.lower())
    if clean_word:  # 过滤空字符串
        clean_bad_words.append(clean_word)

def get_user_mention(mention):
    if mention.startswith('<@') and mention.endswith('>'):
        return client.get_user(int(mention[2:-1]))  # 修复类型转换问题

@client.event
async def on_message(message):
    await client.process_commands(message)
    global automod

    # 跳过机器人自身和其他Bot消息
    if message.author == client.user or message.author.bot:
        return

    if automod:
        # 清洗消息内容:移除干扰字符,转为小写
        cleaned_content = re.sub(r'[.,*/_\- ]', '', message.content.lower())
        # 检查清洗后的内容是否包含任意违规词的纯字母序列
        if any(bad_word in cleaned_content for bad_word in clean_bad_words):
            await message.delete()
            await message.channel.send(f"Nuh uh, {message.author.mention} Please do not use bad words.")

    # 自动审核开关命令
    if message.content.startswith("$automod-toggle") and message.author.id in [822685628201697311, 998625868643565640]:
        automod = not automod
        await message.channel.send(f"Automod is now {'enabled' if automod else 'disabled'}")

# 替换为你的机器人Token后启动
# client.run("YOUR_DISCORD_BOT_TOKEN")

关键优化点

  • 正则字符清洗:用re.sub()批量移除指定干扰字符,比手动替换更高效,能覆盖所有目标干扰符号。
  • 双向预处理:同时对违规词和用户消息做纯字母提取,确保匹配不受干扰字符的位置和类型影响,能识别b.a.d、b-a-d、b a d等变形违规词。
  • 修复原代码问题:补上了Bot初始化步骤、修复了用户ID转换的类型错误、删除了重复的intents设置,修正了语法错误。

内容的提问来源于stack exchange,提问作者John Hawking

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 10:20:03