使用preg_match_all提取含História的</h2>后前两段<p>的方法
Hey there! Let me break down how to handle this task with regex—though I should start by saying: using a proper HTML parser is almost always more reliable for this kind of work (HTML is messy, regex wasn't built for it). But since you asked for a regex solution, here's how to make it work:
Step 1: Match the Target <h2> Tag
First, we need to zero in on the <h2> that contains the text "História". The pattern below accounts for any attributes (like class or id) that might be in the <h2> tag:
<h2[^>]*>.*?História.*?</h2>
Let's unpack this:
<h2[^>]*>: Matches the opening<h2>tag, including any attributes (the[^>]*skips everything until the closing>).*?História.*?: Uses a non-greedy match (.*?) to capture the content inside the<h2>without overshooting past other tags</h2>: Matches the closing tag of the heading
Step 2: Capture the Next Two <p> Elements
Now we extend the regex to grab the first two <p> tags that come right after our target <h2>. We'll add capturing groups to extract these paragraphs:
<h2[^>]*>.*?História.*?</h2>\s*(<p[^>]*>.*?</p>)\s*(<p[^>]*>.*?</p>)
Key parts here:
\s*: Matches any whitespace (newlines, spaces, tabs) between the closing</h2>and the first<p>—this handles messy formatting in the HTML(<p[^>]*>.*?</p>): Each capturing group grabs one<p>tag, including any attributes and its content (again using non-greedy matching to avoid grabbing extra content)
Step 3: Restrict to the Specified Div
Since your main content lives in a specific div, wrap the pattern to only search inside that div. For example, if the div has class="main-content", use this:
<div[^>]*class="main-content"[^>]*>.*?<h2[^>]*>.*?História.*?</h2>\s*(<p[^>]*>.*?</p>)\s*(<p[^>]*>.*?</p>)
The .*? between the div opening tag and the <h2> ensures we don't skip past other content in the div.
A Better Alternative: Use an HTML Parser
Regex can fail if your HTML has nested tags inside <p> elements, inconsistent line breaks, or unexpected attributes. For example, in PHP, you could use DOMDocument and XPath for a robust solution:
$doc = new DOMDocument(); @$doc->loadHTML($url); // The @ suppresses warnings about malformed HTML $xpath = new DOMXPath($doc); // Find the <h2> with "História" inside the target div $targetH2 = $xpath->query('//div[@class="main-content"]//h2[contains(text(), "História")]')->item(0); if ($targetH2) { $capturedParagraphs = []; $nextElement = $targetH2->nextSibling; // Grab the next two <p> siblings while ($nextElement && count($capturedParagraphs) < 2) { if ($nextElement->nodeName === 'p') { $capturedParagraphs[] = $doc->saveHTML($nextElement); } $nextElement = $nextElement->nextSibling; } // $capturedParagraphs now holds your two paragraphs! }
This will handle edge cases regex can't, like nested elements or weird formatting.
内容的提问来源于stack exchange,提问作者Gislef

