Lua实现Discord敏感词过滤机器人:求更高效实现方案
Hey there! Your current Discord bot code works solid for exact sensitive word matches, but generating every possible obfuscated variant upfront (like f4rm from farm) isn't the most efficient route—it would bloat memory and miss edge cases. Let's break down smarter alternatives to handle this:
Key Improvements to Your Approach
1. Normalize Message Content First
Start by standardizing common character substitutions users use to bypass filters. This turns obfuscated text back into its base form before checking against your sensitive word list. For example:
- Replace numbers with their letter equivalents (
4→a,3→e,0→o) - Collapse repeated characters (optional, for catching variants like
poooop)
Here's a Lua function for basic normalization:
local function normalizeMessage(content) local normalized = content:lower() -- Replace common number-to-letter swaps normalized = normalized:gsub('4', 'a') :gsub('3', 'e') :gsub('0', 'o') :gsub('1', 'i') :gsub('5', 's') -- Collapse repeated characters (e.g., "poooop" → "poop") normalized = normalized:gsub('([%a])%1+', '%1') return normalized end
2. Use Precompiled Pattern Matching
Instead of splitting messages into words and checking exact matches, use Lua's pattern matching to scan the normalized message for sensitive words (and their variants). Precompile these patterns once when your bot starts to avoid redundant work.
Update your code with this approach:
local BadWords = { 'poop', 'farm' } -- Switch to a list for easier pattern generation local compiledPatterns = {} -- Precompile patterns for each sensitive word once on bot start for _, word in ipairs(BadWords) do -- Build a pattern that matches the word with common substitutions local patternWord = word:gsub('a', '[a4]') :gsub('e', '[e3]') :gsub('o', '[o0]') :gsub('i', '[i1]') :gsub('s', '[s5]') -- Add word boundaries to avoid partial matches (e.g., "poop" won't flag "pooper") table.insert(compiledPatterns, '%f[%a]' .. patternWord .. '%f[%A]') end client:on('messageCreate', function(msg) local normalizedMsg = normalizeMessage(msg.content) for _, pattern in ipairs(compiledPatterns) do local matchedWord = normalizedMsg:match(pattern) if matchedWord then msg:reply({ embed = { fields = { {name = "Bad Word Detected", value = matchedWord, inline = true} }, color = discordia.Color.fromRGB(114, 137, 218).value } }) break -- Stop after first match, remove if you want to catch all instances end end end)
3. Why This Beats Generating All Variants
- Memory Efficiency: Generating every possible variant of a sensitive word (all letter/number swap combinations) would create an enormous list, especially for longer words. Precompiled patterns avoid this entirely.
- Flexibility: Patterns can handle edge cases like extra spaces or symbols without predefining every possible variant.
- Performance: Scanning a single normalized string with precompiled patterns is faster than splitting messages and checking each word against a massive variant list.
Additional Tips
- Tweak Strictness: If you want to catch more obfuscations (like symbols mixed in, e.g.,
p!oop), add symbol-to-letter replacements to your normalization function. - Avoid False Positives: Stick to word boundaries (
%f[%a]/%f[%A]) to ensure you're matching full words, not substrings in legitimate text. - Expand Substitutions: Add more swaps based on obfuscation tactics you see in your server (e.g.,
@→a,$→s).
内容的提问来源于stack exchange,提问作者Albedo

