PHP preg_replace匹配模式:移除class为reference的span标签内内容
<span class="reference"> with Variable data-code Got it, let's fix this for you. The issue with your current regex is that you've hardcoded the data-code="Z22" value, so it can't handle cases where that code changes. Here are two reliable approaches to solve this:
1. Updated Regular Expression (Simple HTML Scenarios)
If your HTML structure is consistent (no weird edge cases like attribute order swapping or escaped quotes), you can tweak the regex to match any data-code value instead of a fixed one.
Modified Code:
$string = 'Some text <span class="reference" data-code="Z22">Data code</span> More text <span class="reference" data-code="X11">Another reference</span>'; // Regex now matches ANY data-code value (using [^"]+ to capture all non-quote characters) $pattern = '|(?<=<span class=\"reference\" data-code=\"[^\"]+\">)(.*?)(?=<\/span>)|'; $replace = '<a href=""> replaced </a>'; // Use preg_replace instead of looping through matches for cleaner code $modifiedString = preg_replace($pattern, $replace, $string); echo $modifiedString;
How It Works:
[^\"]+replaces the hardcodedZ22— this matches any sequence of characters that aren't double quotes, so it works for anydata-codevalue (letters, numbers, mixed characters, etc.).preg_replacehandles all matching spans in one go, so you don't need to loop through matches manually.
2. DOMDocument Approach (Robust for Complex HTML)
Regex can break if your HTML has unexpected variations (like data-code coming before class, or extra attributes in the span). For a more reliable solution, use PHP's built-in DOM parsing tools:
Code Example:
$string = 'Some text <span class="reference" data-code="Z22">Data code</span> More text <span class="reference" data-code="X11">Another reference</span>'; // Initialize DOMDocument and handle potential HTML parsing warnings libxml_use_internal_errors(true); $dom = new DOMDocument(); // Load HTML without adding default doctype/html/body tags $dom->loadHTML(htmlspecialchars_decode($string), LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD); libxml_clear_errors(); // Use XPath to find all span tags with class="reference" $xpath = new DOMXPath($dom); $referenceSpans = $xpath->query('//span[@class="reference"]'); foreach ($referenceSpans as $span) { // Create the replacement <a> tag $replacementLink = $dom->createElement('a'); $replacementLink->setAttribute('href', ''); $replacementLink->nodeValue = ' replaced '; // Clear the span's current content and add the new link $span->nodeValue = ''; $span->appendChild($replacementLink); } // Convert back to escaped HTML string $modifiedString = htmlspecialchars($dom->saveHTML()); echo $modifiedString;
Why This Is Better:
- It doesn't care about attribute order or extra attributes in the span — it reliably targets all spans with
class="reference", regardless of theirdata-codevalue. - Avoids regex pitfalls with malformed or unpredictable HTML.
Choose the first method if your HTML is simple and consistent, and the second if you need to handle more flexible or complex HTML structures.
内容的提问来源于stack exchange,提问作者D. Vasiliev

