PHP:如何替换文章中外链并排除本站链接?遇DOMDocument与正则问题
Hey there! Let's work through your problem step by step. You're trying to rewrite external links (while leaving your own site's links untouched) but hit two snags: a DOMDocument warning about duplicate target attributes, and a regex that replaces all links—including your internal ones. Here's how to fix both:
1. Adjust the Regex to Exclude Internal Links
Your original regex catches every HTTP/HTTPS link, so we need to add a negative lookahead to skip URLs pointing to your domain. Let's assume your site's domain is mydomain.com (swap this for your actual domain):
<?php $html = '1224 <a href="http://www.google.com">google</a> 567 <a href="https://www.mydomain.com/about">About Us</a>'; $tracking_string = 'http://mydomain.com/goto/'; // Add negative lookahead to exclude your domain $html = preg_replace( '#(<a[^>]+href=")(http|https)(?!://(www\.)?mydomain\.com)([^"> ]+)("?[^>]*>)#is', '\1' . $tracking_string . '\2\3\4', $html ); echo $html; ?>
What the regex change does:
(?!://(www\.)?mydomain\.com): This negative lookahead tells the regex, "only match if the URL does NOT start with://www.mydomain.comor://mydomain.com"- The
iflag makes the match case-insensitive, so it catches variations likeMyDomain.comtoo.
2. Fix the DOMDocument Warning & Use It for Robust Link Handling
The Attribute target redefined warning pops up because your HTML has duplicate target attributes on some <a> tags (e.g., <a href="..." target="_blank" target="_self">). Here's how to suppress that warning and properly rewrite only external links:
<?php $html = '1224 <a href="http://www.google.com" target="_blank" target="_self">google</a> 567 <a href="https://www.mydomain.com/contact">Contact</a>'; $tracking_base = 'http://mydomain.com/goto/'; $internal_domain = 'mydomain.com'; // Suppress libxml warnings (for duplicate attributes) libxml_use_internal_errors(true); $dom = new DOMDocument(); // Load HTML with UTF-8 support to avoid encoding issues $dom->loadHTML('<?xml encoding="UTF-8">' . $html); libxml_clear_errors(); // Clear any stored error messages foreach ($dom->getElementsByTagName('a') as $link) { $href = $link->getAttribute('href'); if (empty($href)) continue; // Parse the URL to check if it's external $url_parts = parse_url($href); // Skip relative links (internal) and links to your domain if (!isset($url_parts['host']) || str_contains($url_parts['host'], $internal_domain)) { continue; } // Rewrite the external link (encode to avoid broken characters) $encoded_href = urlencode($href); $link->setAttribute('href', $tracking_base . $encoded_href); } // Output the modified HTML echo $dom->saveHTML(); ?>
Key fixes here:
libxml_use_internal_errors(true): Turns off the annoying duplicate attribute warnings (DOMDocument automatically handles duplicates by keeping the last value).parse_url(): Properly checks if the link's host matches your internal domain, so we skip rewriting internal links (including relative paths like/about).urlencode(): Ensures the original external URL is safely encoded in your tracking link, avoiding broken characters or malformed URLs.
Which Method Should You Choose?
- Regex: Perfect for simple HTML snippets where you know the structure is clean. It's fast and easy to implement.
- DOMDocument: Better for complex HTML (like full pages with nested elements) because it properly parses the DOM instead of relying on string matching, which can break if your HTML has unexpected formatting.
内容的提问来源于stack exchange,提问作者thanh

