You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在脏话替换的preg_replace中添加唯一ID与匹配词收集功能

Solution for Enhancing Profanity Filter with Span IDs and Match Tracking

Got it, let's build out your enhanced profanity filter. The key here is swapping preg_replace for preg_replace_callback—this lets us run custom logic for each match, which is perfect for tracking matches and generating unique IDs.

Here's a complete, working implementation:

<?php
// Your list of profane words (adjust as needed)
$profanityList = ['ass', 'damn', 'hell'];

// Initialize variables to track matches and generate IDs
$matchedProfanity = [];
$spanIdCounter = 1;

// The input text to filter
$inputText = "Oh damn, that's a stupid ass move! Hell, why would you do that ass thing again?";

// Use preg_replace_callback for dynamic processing
$filteredText = preg_replace_callback(
    '/\b(' . implode('|', array_map('preg_quote', $profanityList)) . ')\b/i',
    function($matches) use (&$matchedProfanity, &$spanIdCounter) {
        // 1. Add the matched profanity to our tracking array (preserve original case)
        $matchedProfanity[] = $matches[0];
        
        // 2. Generate the censored string (first letter + dots)
        $censored = $matches[0][0] . str_repeat('.', strlen($matches[0]) - 1);
        
        // 3. Wrap in span with unique ID
        $span = "<span id='profanity-{$spanIdCounter}'>{$censored}</span>";
        
        // Increment counter for next match
        $spanIdCounter++;
        
        return $span;
    },
    $inputText
);

// Output the filtered text
echo "Filtered Text:\n{$filteredText}\n\n";

// Output the tracked profanity list
echo "Matched Profanity (in order):\n";
foreach ($matchedProfanity as $index => $word) {
    echo "- #" . ($index + 1) . ": {$word}\n";
}
?>

Key Details Explained:

  • preg_replace_callback: This replaces each match with the result of our custom callback, letting us handle span ID generation and match tracking in one go—something regular preg_replace can't do easily.
  • Unique Span IDs: We use a reference variable $spanIdCounter that increments with each match, ensuring every censored word gets a unique id (like profanity-1, profanity-2).
  • Match Tracking: The $matchedProfanity array collects every matched profanity (including duplicates) in the exact order they appear. Using a reference variable lets the callback modify the original array outside the function.
  • Censoring Logic: We keep the first character of the profane word and replace the rest with dots. str_repeat ensures the censored string length matches the original word (optional, but makes the output look more natural).
  • Regex Safety: preg_quote escapes special characters in profanity words to avoid regex syntax conflicts, and \b word boundaries prevent partial matches (e.g., not triggering on "ass" in "class").

Example Output:

Filtered Text:
Oh <span id='profanity-1'>d...</span>, that's a stupid <span id='profanity-2'>a..</span> move! <span id='profanity-3'>H...</span>, why would you do that <span id='profanity-4'>a..</span> thing again?

Matched Profanity (in order):
- #1: damn
- #2: ass
- #3: Hell
- #4: ass

This implementation covers both of your requested features while keeping the original censoring behavior intact. Feel free to tweak the profanity list or censoring logic to fit your specific use case!

内容的提问来源于stack exchange,提问作者Andres SK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:48:48