如何实现基于不良词数组的文本过滤与替换功能?
Got it, let's break this down step by step—since you're focused on learning the implementation pattern rather than just building a one-off feature, I’ll walk through the core ideas first, then show you practical examples in common languages, plus an optimized approach for larger datasets.
核心逻辑拆解
Before diving into code, let's outline the key pieces you need to cover:
- Matching Strategy: Choose how to efficiently find bad words in the input text. Simple regex works for small bad word lists, while more advanced structures like Trie trees (dictionary trees) are better for large datasets.
- Replacement Logic: Swap out any matched bad words with
[removed]. - Edge Case Handling: Avoid false positives (e.g., don't replace "bad" in "badly") by using word boundaries or longest-match rules.
Basic Implementation: Regular Expressions
Regex is a straightforward starting point for small to medium bad word lists. It’s easy to write and understand, which makes it great for learning the core flow.
Python Example
import re def filter_bad_words(message: str, bad_words: list) -> str: # Escape special characters in bad words to avoid regex syntax conflicts escaped_words = [re.escape(word) for word in bad_words] # Build a regex pattern with word boundaries to prevent partial matches pattern = rf'\b({"|".join(escaped_words)})\b' # Replace matches with [removed], ignoring case return re.sub(pattern, '[removed]', message, flags=re.IGNORECASE) # Test it out bad_words = ["垃圾", "笨蛋", "stupid"] user_message = "你这个笨蛋,说什么垃圾话呢,Stupid guy!" print(filter_bad_words(user_message, bad_words)) # Output: 你这个[removed],说什么[removed]话呢,[removed] guy!
JavaScript Example
function filterBadWords(message, badWords) { // Escape regex special characters in each bad word const escapedWords = badWords.map(word => word.replace(/[.*+?^${}()|[\]\\]/g, '\\$&') ); // Create regex with word boundaries and case-insensitive matching const pattern = new RegExp(`\\b(${escapedWords.join('|')})\\b`, 'gi'); // Replace matches return message.replace(pattern, '[removed]'); } // Test const badWords = ["垃圾", "笨蛋", "stupid"]; const userMessage = "你这个笨蛋,说什么垃圾话呢,Stupid guy!"; console.log(filterBadWords(userMessage, badWords)); // Output: 你这个[removed],说什么[removed]话呢,[removed] guy!
Optimized Approach: Trie Tree (For Large Bad Word Lists)
If you’re dealing with hundreds or thousands of bad words, regex performance starts to drop. A Trie tree lets you traverse the input text once to find all matches, making it much more efficient.
Python Trie Tree Implementation
class TrieNode: def __init__(self): self.children = {} self.is_end_of_word = False class BadWordFilter: def __init__(self, bad_words): self.root = TrieNode() self._build_trie(bad_words) def _build_trie(self, bad_words): # Build the trie from the bad word list (case-insensitive) for word in bad_words: current_node = self.root for char in word.lower(): if char not in current_node.children: current_node.children[char] = TrieNode() current_node = current_node.children[char] current_node.is_end_of_word = True def filter(self, message): message_lower = message.lower() result = [] index = 0 message_length = len(message) while index < message_length: current_node = self.root match_end_index = index longest_match_length = 0 # Look for the longest possible bad word match starting at current index while match_end_index < message_length and message_lower[match_end_index] in current_node.children: current_node = current_node.children[message_lower[match_end_index]] match_end_index += 1 if current_node.is_end_of_word: longest_match_length = match_end_index - index if longest_match_length > 0: result.append('[removed]') index = match_end_index else: result.append(message[index]) index += 1 return ''.join(result) # Test filter = BadWordFilter(["垃圾", "笨蛋", "stupid"]) user_message = "你这个笨蛋,说什么垃圾话呢,Stupid guy!" print(filter.filter(user_message)) # Output: 你这个[removed],说什么[removed]话呢,[removed] guy!
Key Considerations
- Case Insensitivity: Always normalize text (e.g., convert to lowercase) to catch variants like "Stupid" and "stupid".
- Special Characters: Use regex escaping or Trie tree handling to account for bad words with special characters (like "fu*k").
- False Positives: Word boundaries (
\bin regex) or longest-match logic (in Trie) prevent accidental replacement of partial matches.
内容的提问来源于stack exchange,提问作者Gregory N. M.

