You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用preg_match_all提取含História的</h2>后前两段<p>的方法

How to Target the "História" <h2> and Capture Its Next Two <p> Tags with Regex

Hey there! Let me break down how to handle this task with regex—though I should start by saying: using a proper HTML parser is almost always more reliable for this kind of work (HTML is messy, regex wasn't built for it). But since you asked for a regex solution, here's how to make it work:

Step 1: Match the Target <h2> Tag

First, we need to zero in on the <h2> that contains the text "História". The pattern below accounts for any attributes (like class or id) that might be in the <h2> tag:

<h2[^>]*>.*?História.*?</h2>

Let's unpack this:

  • <h2[^>]*>: Matches the opening <h2> tag, including any attributes (the [^>]* skips everything until the closing >)
  • .*?História.*?: Uses a non-greedy match (.*?) to capture the content inside the <h2> without overshooting past other tags
  • </h2>: Matches the closing tag of the heading

Step 2: Capture the Next Two <p> Elements

Now we extend the regex to grab the first two <p> tags that come right after our target <h2>. We'll add capturing groups to extract these paragraphs:

<h2[^>]*>.*?História.*?</h2>\s*(<p[^>]*>.*?</p>)\s*(<p[^>]*>.*?</p>)

Key parts here:

  • \s*: Matches any whitespace (newlines, spaces, tabs) between the closing </h2> and the first <p>—this handles messy formatting in the HTML
  • (<p[^>]*>.*?</p>): Each capturing group grabs one <p> tag, including any attributes and its content (again using non-greedy matching to avoid grabbing extra content)

Step 3: Restrict to the Specified Div

Since your main content lives in a specific div, wrap the pattern to only search inside that div. For example, if the div has class="main-content", use this:

<div[^>]*class="main-content"[^>]*>.*?<h2[^>]*>.*?História.*?</h2>\s*(<p[^>]*>.*?</p>)\s*(<p[^>]*>.*?</p>)

The .*? between the div opening tag and the <h2> ensures we don't skip past other content in the div.

A Better Alternative: Use an HTML Parser

Regex can fail if your HTML has nested tags inside <p> elements, inconsistent line breaks, or unexpected attributes. For example, in PHP, you could use DOMDocument and XPath for a robust solution:

$doc = new DOMDocument();
@$doc->loadHTML($url); // The @ suppresses warnings about malformed HTML
$xpath = new DOMXPath($doc);

// Find the <h2> with "História" inside the target div
$targetH2 = $xpath->query('//div[@class="main-content"]//h2[contains(text(), "História")]')->item(0);

if ($targetH2) {
    $capturedParagraphs = [];
    $nextElement = $targetH2->nextSibling;
    
    // Grab the next two <p> siblings
    while ($nextElement && count($capturedParagraphs) < 2) {
        if ($nextElement->nodeName === 'p') {
            $capturedParagraphs[] = $doc->saveHTML($nextElement);
        }
        $nextElement = $nextElement->nextSibling;
    }
    
    // $capturedParagraphs now holds your two paragraphs!
}

This will handle edge cases regex can't, like nested elements or weird formatting.

内容的提问来源于stack exchange,提问作者Gislef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:03:48