PHP中将HTML链接拆分为含标题与链接的数组方法问询
Great question! Let's break down two solid approaches to convert that escaped HTML into the array of link data you're after. The first method is the most reliable (since HTML can be tricky with regex), and the second uses regex if you're dealing with strictly consistent markup.
方法1:使用DOMDocument解析(推荐)
When working with HTML, using a proper DOM parser is always better than regex—it handles edge cases like nested tags, attribute order changes, or malformed markup way more gracefully. Here's how to do it in PHP:
步骤:
- First, convert the escaped HTML entities (
<,>) back to actual HTML tags withhtmlspecialchars_decode(). - Load the cleaned HTML into
DOMDocument. - Grab all
<a>tags usinggetElementsByTagName(). - Loop through each tag to extract the
hrefattribute and link text, then build your array.
代码示例:
// Your original escaped HTML string $escapedHtml = '<a href="/newsitems">News</a> <a href="/news/roman-catapults/16465">Roman Catapults</a> <a href="/news/year-3-roman-experience/13835">Year 3 Roman Experience</a> <a href="/news/year-3-dewa-roman-experience/15746">Year 3 Dewa Roman Experience</a> <a href="/news/science-week-day-1/15423">Science Week</a><a href="/news/world-book-day/15104">World Book Day</a> <a href="/news/year-6-trip-to-the-lion-salt-works/15762">Year 6 trip to the Lion Salt Works</a><a href="/news/learning-logs/13839">Learning Logs</a> <a href="/news/working-together/13838">Working Together</a> <a href="/news/learning-logs/13837">Learning Logs</a> <a href="/news/year-2-curriculum-map-for-autumn-2/13377">Year 2 Curriculum Map for Autumn 2</a>'; // Convert escaped entities back to real HTML tags $html = htmlspecialchars_decode($escapedHtml); // Initialize DOMDocument and suppress minor HTML warnings $dom = new DOMDocument(); libxml_use_internal_errors(true); // Ignore any parsing warnings $dom->loadHTML($html); libxml_clear_errors(); // Clear the error buffer // Get all <a> elements $linkElements = $dom->getElementsByTagName('a'); $linkArray = []; // Loop through each link and extract data foreach ($linkElements as $link) { $href = $link->getAttribute('href'); $title = trim($link->nodeValue); // Trim extra whitespace from the text // Only add to the array if both href and title exist if (!empty($href) && !empty($title)) { $linkArray[] = [ 'title' => $title, 'link' => $href ]; } } // Output the result print_r($linkArray);
为什么推荐这个方法?
- It handles messy HTML: If the
<a>tags ever get extra attributes, nested elements (like<span>inside), or different formatting, this will still work. - It's maintainable: You don't have to tweak regex patterns every time the HTML structure changes.
方法2:使用正则表达式(适合固定结构)
If you're 100% sure the HTML will never change (no nested tags, consistent attribute order, double quotes for href), regex can work. Here's how:
步骤:
- Again, convert the escaped HTML to real tags first.
- Use
preg_match_all()to capture thehrefvalue and link text from each<a>tag. - Map the captured groups into your desired array structure.
代码示例:
$escapedHtml = '<a href="/newsitems">News</a> <a href="/news/roman-catapults/16465">Roman Catapults</a> <a href="/news/year-3-roman-experience/13835">Year 3 Roman Experience</a> <a href="/news/year-3-dewa-roman-experience/15746">Year 3 Dewa Roman Experience</a> <a href="/news/science-week-day-1/15423">Science Week</a><a href="/news/world-book-day/15104">World Book Day</a> <a href="/news/year-6-trip-to-the-lion-salt-works/15762">Year 6 trip to the Lion Salt Works</a><a href="/news/learning-logs/13839">Learning Logs</a> <a href="/news/working-together/13838">Working Together</a> <a href="/news/learning-logs/13837">Learning Logs</a> <a href="/news/year-2-curriculum-map-for-autumn-2/13377">Year 2 Curriculum Map for Autumn 2</a>'; $html = htmlspecialchars_decode($escapedHtml); // Regex pattern to match <a> tags: captures href value and link text $pattern = '/<a href="([^"]+)">([^<]+)<\/a>/i'; // Run the match preg_match_all($pattern, $html, $matches); $linkArray = []; // Pair up the captured hrefs and titles foreach ($matches[1] as $index => $href) { $title = trim($matches[2][$index]); if (!empty($href) && !empty($title)) { $linkArray[] = [ 'title' => $title, 'link' => $href ]; } } print_r($linkArray);
注意事项:
- This regex assumes:
hrefuses double quotes (won't work if it's single quotes likehref='...').- No nested tags inside the
<a>(e.g.,<a href="..."><b>Text</b></a>will break this). - The link text doesn't contain
<characters (which it shouldn't in valid HTML, but just in case).
- If the HTML ever changes, you'll have to update the regex pattern.
Final Recommendation
Stick with the DOMDocument method for most cases—it's the industry standard for parsing HTML and will save you headaches down the line. Use regex only if you're dealing with a completely static, unchanging HTML snippet.
内容的提问来源于stack exchange,提问作者YaBCK

