You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现基于不良词数组的文本过滤与替换功能?

Got it, let's break this down step by step—since you're focused on learning the implementation pattern rather than just building a one-off feature, I’ll walk through the core ideas first, then show you practical examples in common languages, plus an optimized approach for larger datasets.

不良词过滤与替换:实现思路与代码示例

核心逻辑拆解

Before diving into code, let's outline the key pieces you need to cover:

  • Matching Strategy: Choose how to efficiently find bad words in the input text. Simple regex works for small bad word lists, while more advanced structures like Trie trees (dictionary trees) are better for large datasets.
  • Replacement Logic: Swap out any matched bad words with [removed].
  • Edge Case Handling: Avoid false positives (e.g., don't replace "bad" in "badly") by using word boundaries or longest-match rules.

Basic Implementation: Regular Expressions

Regex is a straightforward starting point for small to medium bad word lists. It’s easy to write and understand, which makes it great for learning the core flow.

Python Example

import re

def filter_bad_words(message: str, bad_words: list) -> str:
    # Escape special characters in bad words to avoid regex syntax conflicts
    escaped_words = [re.escape(word) for word in bad_words]
    # Build a regex pattern with word boundaries to prevent partial matches
    pattern = rf'\b({"|".join(escaped_words)})\b'
    # Replace matches with [removed], ignoring case
    return re.sub(pattern, '[removed]', message, flags=re.IGNORECASE)

# Test it out
bad_words = ["垃圾", "笨蛋", "stupid"]
user_message = "你这个笨蛋,说什么垃圾话呢,Stupid guy!"
print(filter_bad_words(user_message, bad_words))
# Output: 你这个[removed],说什么[removed]话呢,[removed] guy!

JavaScript Example

function filterBadWords(message, badWords) {
    // Escape regex special characters in each bad word
    const escapedWords = badWords.map(word => 
        word.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
    );
    // Create regex with word boundaries and case-insensitive matching
    const pattern = new RegExp(`\\b(${escapedWords.join('|')})\\b`, 'gi');
    // Replace matches
    return message.replace(pattern, '[removed]');
}

// Test
const badWords = ["垃圾", "笨蛋", "stupid"];
const userMessage = "你这个笨蛋,说什么垃圾话呢,Stupid guy!";
console.log(filterBadWords(userMessage, badWords));
// Output: 你这个[removed],说什么[removed]话呢,[removed] guy!

Optimized Approach: Trie Tree (For Large Bad Word Lists)

If you’re dealing with hundreds or thousands of bad words, regex performance starts to drop. A Trie tree lets you traverse the input text once to find all matches, making it much more efficient.

Python Trie Tree Implementation

class TrieNode:
    def __init__(self):
        self.children = {}
        self.is_end_of_word = False

class BadWordFilter:
    def __init__(self, bad_words):
        self.root = TrieNode()
        self._build_trie(bad_words)
    
    def _build_trie(self, bad_words):
        # Build the trie from the bad word list (case-insensitive)
        for word in bad_words:
            current_node = self.root
            for char in word.lower():
                if char not in current_node.children:
                    current_node.children[char] = TrieNode()
                current_node = current_node.children[char]
            current_node.is_end_of_word = True
    
    def filter(self, message):
        message_lower = message.lower()
        result = []
        index = 0
        message_length = len(message)
        
        while index < message_length:
            current_node = self.root
            match_end_index = index
            longest_match_length = 0
            
            # Look for the longest possible bad word match starting at current index
            while match_end_index < message_length and message_lower[match_end_index] in current_node.children:
                current_node = current_node.children[message_lower[match_end_index]]
                match_end_index += 1
                if current_node.is_end_of_word:
                    longest_match_length = match_end_index - index
            
            if longest_match_length > 0:
                result.append('[removed]')
                index = match_end_index
            else:
                result.append(message[index])
                index += 1
        
        return ''.join(result)

# Test
filter = BadWordFilter(["垃圾", "笨蛋", "stupid"])
user_message = "你这个笨蛋,说什么垃圾话呢,Stupid guy!"
print(filter.filter(user_message))
# Output: 你这个[removed],说什么[removed]话呢,[removed] guy!

Key Considerations

  • Case Insensitivity: Always normalize text (e.g., convert to lowercase) to catch variants like "Stupid" and "stupid".
  • Special Characters: Use regex escaping or Trie tree handling to account for bad words with special characters (like "fu*k").
  • False Positives: Word boundaries (\b in regex) or longest-match logic (in Trie) prevent accidental replacement of partial matches.

内容的提问来源于stack exchange,提问作者Gregory N. M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 07:47:41