You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP中将HTML链接拆分为含标题与链接的数组方法问询

Great question! Let's break down two solid approaches to convert that escaped HTML into the array of link data you're after. The first method is the most reliable (since HTML can be tricky with regex), and the second uses regex if you're dealing with strictly consistent markup.

方法1:使用DOMDocument解析(推荐)

When working with HTML, using a proper DOM parser is always better than regex—it handles edge cases like nested tags, attribute order changes, or malformed markup way more gracefully. Here's how to do it in PHP:

步骤:

  1. First, convert the escaped HTML entities (<, >) back to actual HTML tags with htmlspecialchars_decode().
  2. Load the cleaned HTML into DOMDocument.
  3. Grab all <a> tags using getElementsByTagName().
  4. Loop through each tag to extract the href attribute and link text, then build your array.

代码示例:

// Your original escaped HTML string
$escapedHtml = '&lt;a href="/newsitems"&gt;News&lt;/a&gt; &lt;a href="/news/roman-catapults/16465"&gt;Roman Catapults&lt;/a&gt; &lt;a href="/news/year-3-roman-experience/13835"&gt;Year 3 Roman Experience&lt;/a&gt; &lt;a href="/news/year-3-dewa-roman-experience/15746"&gt;Year 3 Dewa Roman Experience&lt;/a&gt; &lt;a href="/news/science-week-day-1/15423"&gt;Science Week&lt;/a&gt;&lt;a href="/news/world-book-day/15104"&gt;World Book Day&lt;/a&gt; &lt;a href="/news/year-6-trip-to-the-lion-salt-works/15762"&gt;Year 6 trip to the Lion Salt Works&lt;/a&gt;&lt;a href="/news/learning-logs/13839"&gt;Learning Logs&lt;/a&gt; &lt;a href="/news/working-together/13838"&gt;Working Together&lt;/a&gt; &lt;a href="/news/learning-logs/13837"&gt;Learning Logs&lt;/a&gt; &lt;a href="/news/year-2-curriculum-map-for-autumn-2/13377"&gt;Year 2 Curriculum Map for Autumn 2&lt;/a&gt;';

// Convert escaped entities back to real HTML tags
$html = htmlspecialchars_decode($escapedHtml);

// Initialize DOMDocument and suppress minor HTML warnings
$dom = new DOMDocument();
libxml_use_internal_errors(true); // Ignore any parsing warnings
$dom->loadHTML($html);
libxml_clear_errors(); // Clear the error buffer

// Get all <a> elements
$linkElements = $dom->getElementsByTagName('a');

$linkArray = [];

// Loop through each link and extract data
foreach ($linkElements as $link) {
    $href = $link->getAttribute('href');
    $title = trim($link->nodeValue); // Trim extra whitespace from the text

    // Only add to the array if both href and title exist
    if (!empty($href) && !empty($title)) {
        $linkArray[] = [
            'title' => $title,
            'link' => $href
        ];
    }
}

// Output the result
print_r($linkArray);

为什么推荐这个方法?

  • It handles messy HTML: If the <a> tags ever get extra attributes, nested elements (like <span> inside), or different formatting, this will still work.
  • It's maintainable: You don't have to tweak regex patterns every time the HTML structure changes.

方法2:使用正则表达式(适合固定结构)

If you're 100% sure the HTML will never change (no nested tags, consistent attribute order, double quotes for href), regex can work. Here's how:

步骤:

  1. Again, convert the escaped HTML to real tags first.
  2. Use preg_match_all() to capture the href value and link text from each <a> tag.
  3. Map the captured groups into your desired array structure.

代码示例:

$escapedHtml = '&lt;a href="/newsitems"&gt;News&lt;/a&gt; &lt;a href="/news/roman-catapults/16465"&gt;Roman Catapults&lt;/a&gt; &lt;a href="/news/year-3-roman-experience/13835"&gt;Year 3 Roman Experience&lt;/a&gt; &lt;a href="/news/year-3-dewa-roman-experience/15746"&gt;Year 3 Dewa Roman Experience&lt;/a&gt; &lt;a href="/news/science-week-day-1/15423"&gt;Science Week&lt;/a&gt;&lt;a href="/news/world-book-day/15104"&gt;World Book Day&lt;/a&gt; &lt;a href="/news/year-6-trip-to-the-lion-salt-works/15762"&gt;Year 6 trip to the Lion Salt Works&lt;/a&gt;&lt;a href="/news/learning-logs/13839"&gt;Learning Logs&lt;/a&gt; &lt;a href="/news/working-together/13838"&gt;Working Together&lt;/a&gt; &lt;a href="/news/learning-logs/13837"&gt;Learning Logs&lt;/a&gt; &lt;a href="/news/year-2-curriculum-map-for-autumn-2/13377"&gt;Year 2 Curriculum Map for Autumn 2&lt;/a&gt;';
$html = htmlspecialchars_decode($escapedHtml);

// Regex pattern to match <a> tags: captures href value and link text
$pattern = '/<a href="([^"]+)">([^<]+)<\/a>/i';

// Run the match
preg_match_all($pattern, $html, $matches);

$linkArray = [];

// Pair up the captured hrefs and titles
foreach ($matches[1] as $index => $href) {
    $title = trim($matches[2][$index]);
    if (!empty($href) && !empty($title)) {
        $linkArray[] = [
            'title' => $title,
            'link' => $href
        ];
    }
}

print_r($linkArray);

注意事项:

  • This regex assumes:
    • href uses double quotes (won't work if it's single quotes like href='...').
    • No nested tags inside the <a> (e.g., <a href="..."><b>Text</b></a> will break this).
    • The link text doesn't contain < characters (which it shouldn't in valid HTML, but just in case).
  • If the HTML ever changes, you'll have to update the regex pattern.

Final Recommendation

Stick with the DOMDocument method for most cases—it's the industry standard for parsing HTML and will save you headaches down the line. Use regex only if you're dealing with a completely static, unchanging HTML snippet.

内容的提问来源于stack exchange,提问作者YaBCK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:50:51