You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP:如何替换文章中外链并排除本站链接?遇DOMDocument与正则问题

Hey there! Let's work through your problem step by step. You're trying to rewrite external links (while leaving your own site's links untouched) but hit two snags: a DOMDocument warning about duplicate target attributes, and a regex that replaces all links—including your internal ones. Here's how to fix both:

Your original regex catches every HTTP/HTTPS link, so we need to add a negative lookahead to skip URLs pointing to your domain. Let's assume your site's domain is mydomain.com (swap this for your actual domain):

<?php
$html = '1224 <a href="http://www.google.com">google</a> 567 <a href="https://www.mydomain.com/about">About Us</a>';
$tracking_string = 'http://mydomain.com/goto/';
// Add negative lookahead to exclude your domain
$html = preg_replace(
    '#(<a[^>]+href=")(http|https)(?!://(www\.)?mydomain\.com)([^"> ]+)("?[^>]*>)#is',
    '\1' . $tracking_string . '\2\3\4',
    $html
);
echo $html;
?>

What the regex change does:

  • (?!://(www\.)?mydomain\.com): This negative lookahead tells the regex, "only match if the URL does NOT start with ://www.mydomain.com or ://mydomain.com"
  • The i flag makes the match case-insensitive, so it catches variations like MyDomain.com too.

The Attribute target redefined warning pops up because your HTML has duplicate target attributes on some <a> tags (e.g., <a href="..." target="_blank" target="_self">). Here's how to suppress that warning and properly rewrite only external links:

<?php
$html = '1224 <a href="http://www.google.com" target="_blank" target="_self">google</a> 567 <a href="https://www.mydomain.com/contact">Contact</a>';
$tracking_base = 'http://mydomain.com/goto/';
$internal_domain = 'mydomain.com';

// Suppress libxml warnings (for duplicate attributes)
libxml_use_internal_errors(true);

$dom = new DOMDocument();
// Load HTML with UTF-8 support to avoid encoding issues
$dom->loadHTML('<?xml encoding="UTF-8">' . $html);
libxml_clear_errors(); // Clear any stored error messages

foreach ($dom->getElementsByTagName('a') as $link) {
    $href = $link->getAttribute('href');
    if (empty($href)) continue;

    // Parse the URL to check if it's external
    $url_parts = parse_url($href);
    // Skip relative links (internal) and links to your domain
    if (!isset($url_parts['host']) || str_contains($url_parts['host'], $internal_domain)) {
        continue;
    }

    // Rewrite the external link (encode to avoid broken characters)
    $encoded_href = urlencode($href);
    $link->setAttribute('href', $tracking_base . $encoded_href);
}

// Output the modified HTML
echo $dom->saveHTML();
?>

Key fixes here:

  • libxml_use_internal_errors(true): Turns off the annoying duplicate attribute warnings (DOMDocument automatically handles duplicates by keeping the last value).
  • parse_url(): Properly checks if the link's host matches your internal domain, so we skip rewriting internal links (including relative paths like /about).
  • urlencode(): Ensures the original external URL is safely encoded in your tracking link, avoiding broken characters or malformed URLs.

Which Method Should You Choose?

  • Regex: Perfect for simple HTML snippets where you know the structure is clean. It's fast and easy to implement.
  • DOMDocument: Better for complex HTML (like full pages with nested elements) because it properly parses the DOM instead of relying on string matching, which can break if your HTML has unexpected formatting.

内容的提问来源于stack exchange,提问作者thanh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:15:13