如何使用DOMXPath获取<h2>标签后div内的<p>内容?
How to Access
Content Inside a
That Follows an 1. Grab the first
with DOMXPath
Hey there! I’ve hit this exact XPath roadblock before—let’s work through how to fix it. First, let’s assume your HTML structure looks like this (super common for this scenario):
<h2>My Awesome Heading</h2> <div class="section-content"> <p>This is the paragraph we need to grab!</p> <p>Maybe there's a second one too!</p> </div>
The most likely issue is that you’re using child-node syntax (/div) instead of targeting sibling nodes. Here’s how to fix it:
1. Grab the first immediately after a specific
If the
is the direct next sibling of your , use following-sibling::div[1] to target it, then drill down to the
tags:
$dom = new DOMDocument();
libxml_use_internal_errors(true); // Suppress messy HTML parsing warnings
$dom->loadHTML($yourHtmlString);
$xpath = new DOMXPath($dom);
// Target the <h2> with specific text, then its next div, then all <p>s
$targetHeading = "My Awesome Heading";
$paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following-sibling::div[1]/p");
// Loop through results if there are multiple paragraphs
foreach ($paragraphs as $p) {
echo trim($p->nodeValue) . "\n";
}
2. Grab a that comes after the (not necessarily immediately)
If there are other elements between the
and , use following::div[1] instead—it finds the first anywhere after the :
$paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following::div[1]/p");
Common Mistakes to Avoid
- Using
/div instead of sibling syntax: /div looks for child nodes of the , but your is a sibling, not a child. That’s almost certainly why your initial queries failed!
- Forgetting
normalize-space(): If your has extra whitespace (like newlines or tabs), text() won’t match exactly. normalize-space() cleans that up.
- Not specifying
[1]: If there are multiple s after the , omitting [1] will grab all of them—adjust based on your needs.
If your HTML structure is a bit different (like nested elements), share the exact markup and I can tweak the XPath further—but these patterns work for most common cases.
内容的提问来源于stack exchange,提问作者Gislef
If the
is the direct next sibling of your , use
2. Grab a
, use following-sibling::div[1] to target it, then drill down to the
tags:
$dom = new DOMDocument(); libxml_use_internal_errors(true); // Suppress messy HTML parsing warnings $dom->loadHTML($yourHtmlString); $xpath = new DOMXPath($dom); // Target the <h2> with specific text, then its next div, then all <p>s $targetHeading = "My Awesome Heading"; $paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following-sibling::div[1]/p"); // Loop through results if there are multiple paragraphs foreach ($paragraphs as $p) { echo trim($p->nodeValue) . "\n"; }
2. Grab a that comes after the (not necessarily immediately)
If there are other elements between the
and , use following::div[1] instead—it finds the first anywhere after the :
$paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following::div[1]/p");
Common Mistakes to Avoid
- Using
/div instead of sibling syntax: /div looks for child nodes of the , but your is a sibling, not a child. That’s almost certainly why your initial queries failed!
- Forgetting
normalize-space(): If your has extra whitespace (like newlines or tabs), text() won’t match exactly. normalize-space() cleans that up.
- Not specifying
[1]: If there are multiple s after the , omitting [1] will grab all of them—adjust based on your needs.
If your HTML structure is a bit different (like nested elements), share the exact markup and I can tweak the XPath further—but these patterns work for most common cases.
内容的提问来源于stack exchange,提问作者Gislef
(not necessarily immediately)
If there are other elements between the
and , use following::div[1] instead—it finds the first anywhere after the :
$paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following::div[1]/p");
Common Mistakes to Avoid
- Using
/div instead of sibling syntax: /div looks for child nodes of the , but your is a sibling, not a child. That’s almost certainly why your initial queries failed!
- Forgetting
normalize-space(): If your has extra whitespace (like newlines or tabs), text() won’t match exactly. normalize-space() cleans that up.
- Not specifying
[1]: If there are multiple s after the , omitting [1] will grab all of them—adjust based on your needs.
If your HTML structure is a bit different (like nested elements), share the exact markup and I can tweak the XPath further—but these patterns work for most common cases.
内容的提问来源于stack exchange,提问作者Gislef
following::div[1] instead—it finds the first anywhere after the
:
$paragraphs = $xpath->query("//h2[normalize-space(text())='$targetHeading']/following::div[1]/p");
Common Mistakes to Avoid
- Using
/divinstead of sibling syntax:/divlooks for child nodes of the, but your
is a sibling, not a child. That’s almost certainly why your initial queries failed! - Forgetting
normalize-space(): If yourhas extra whitespace (like newlines or tabs),
text()won’t match exactly.normalize-space()cleans that up. - Not specifying
[1]: If there are multiples after the, omitting
[1]will grab all of them—adjust based on your needs.
If your HTML structure is a bit different (like nested elements), share the exact markup and I can tweak the XPath further—but these patterns work for most common cases.
内容的提问来源于stack exchange,提问作者Gislef
相关产品推荐
相关产品推荐

