如何编写正则表达式匹配黑名单词汇的各类变形?
Hey there! Let's fix this regex problem you're having. The issue with your current patterns is they don't account for the separators (_,-,., etc.) between the letters of your blacklisted words. Let's break down how to build a regex that catches all those deformed versions.
Why Your Current Regexes Fall Short
- Your first pattern
\W*((?i)ask(?-i))\W*only matches the exact word "ask" (case-insensitive) surrounded by non-word characters—it won't catcha_ksoras-kbecause the separators are between the letters, not around the whole word. - The second pattern
^((?!ask)[\s\S])*$is a negative lookahead that blocks any input containing the exact string "ask", but again, it ignores any separators between the letters, soa.s.kwould slip right through.
The Solution: Regex That Allows Separators Between Letters
The key idea is to split each blacklisted word into individual characters, and between each character, allow zero or more of your target separators (_, -, ., etc.). We'll also make the whole pattern case-insensitive so it catches Ask, ASK, aSk, etc.
Example Regex for Your Blacklisted Words
For the words ask, stackoverflow, and hello, here's a regex that will match all their deformed versions:
$blacklistRegex = '/(?i)(?:a[_\-.]*s[_\-.]*k|s[_\-.]*t[_\-.]*a[_\-.]*c[_\-.]*k[_\-.]*o[_\-.]*v[_\-.]*e[_\-.]*r[_\-.]*f[_\-.]*l[_\-.]*o[_\-.]*w|h[_\-.]*e[_\-.]*l[_\-.]*l[_\-.]*o)/';
Let's Break It Down
(?i): Enables case-insensitive matching so we don't have to worry about uppercase/lowercase variations.(?:...): A non-capturing group to wrap our alternated patterns (keeps the regex efficient).- For each word: Split into characters, and insert
[_\-.]*between each pair. This means "zero or more of underscore, hyphen, or dot" can appear between any two letters of the blacklisted word.a[_\-.]*s[_\-.]*k: Matchesask,a_sk,as-k,a.s.k,A-S-K, etc.- The same logic applies to
stackoverflowandhello.
Adjusting Separators
If you want to allow more separators (like spaces, commas, or other symbols), just add them to the character set [_\-.]. For example, to include spaces and commas: [_\-.,\s].
Using It in PHP
To check if a form input contains any of these deformed blacklisted words, you can use preg_match:
$input = $_POST['user_input']; $blacklistRegex = '/(?i)(?:a[_\-.]*s[_\-.]*k|s[_\-.]*t[_\-.]*a[_\-.]*c[_\-.]*k[_\-.]*o[_\-.]*v[_\-.]*e[_\-.]*r[_\-.]*f[_\-.]*l[_\-.]*o[_\-.]*w|h[_\-.]*e[_\-.]*l[_\-.]*l[_\-.]*o)/'; if (preg_match($blacklistRegex, $input)) { // Input contains a blacklisted word (deformed or not) echo "Invalid input: Contains restricted content."; } else { // Input is clean echo "Input is valid."; }
Automating Regex Generation
If you have a long list of blacklisted words, manually writing the regex for each one is tedious. You can generate the regex programmatically in PHP:
$blacklistedWords = ['ask', 'stackoverflow', 'hello']; $separatorPattern = '[_\-.]*'; $regexParts = []; foreach ($blacklistedWords as $word) { // Split the word into individual characters $chars = str_split($word); // Join with the separator pattern $regexParts[] = implode($separatorPattern, $chars); } // Combine into a single regex $blacklistRegex = '/(?i)(?:' . implode('|', $regexParts) . ')/';
This will dynamically build the regex for any number of blacklisted words, saving you time.
Bonus: Avoid False Positives
If you want to ensure the match isn't part of a longer alphanumeric word (e.g., don't match "task" as a deformed "ask"), you can add lookarounds to check that the characters before and after aren't letters/numbers:
$blacklistRegex = '/(?i)(?:(?<![\p{L}\p{N}])a[_\-.]*s[_\-.]*k(?![\p{L}\p{N}])|(?<![\p{L}\p{N}])s[_\-.]*t[_\-.]*a[_\-.]*c[_\-.]*k[_\-.]*o[_\-.]*v[_\-.]*e[_\-.]*r[_\-.]*f[_\-.]*l[_\-.]*o[_\-.]*w(?![\p{L}\p{N}])|(?<![\p{L}\p{N}])h[_\-.]*e[_\-.]*l[_\-.]*l[_\-.]*o(?![\p{L}\p{N}]))/';
The (?<![\p{L}\p{N}]) is a negative lookbehind (no letter/number before), and (?![\p{L}\p{N}]) is a negative lookahead (no letter/number after).
That should cover all the deformed versions you mentioned. Let me know if you need further tweaks!
内容的提问来源于stack exchange,提问作者DolDurma

