如何编写XPath获取XML中指定h4标题后的列表项内容?
Your initial XPath didn't work because the <ul> is a sibling of the <h4>, not a child. The <h4> and <ul> are both nested inside the same parent <div>, so we need to traverse to the following sibling <ul> instead of looking for a direct child.
Here are the precise, position-agnostic XPath expressions for each heading:
Source:
//h4[normalize-space(text())="Source"]/following-sibling::ul[1]/li/text()normalize-space(text())handles accidental whitespace in the h4 text (like leading/trailing spaces)following-sibling::ul[1]targets the very first<ul>immediately after the matching<h4>, ensuring we grab the correct associated list even if other elements are nearby
Language:
//h4[normalize-space(text())="Language"]/following-sibling::ul[1]/li/text()- This captures the text inside the
<a>tag (e.g., "Swedish" in your example) becausetext()retrieves all text nodes within the<li>
- This captures the text inside the
Publishers:
//h4[normalize-space(text())="Publishers"]/following-sibling::ul[1]/li/text()Record ID:
//h4[normalize-space(text())="Record ID"]/following-sibling::ul[1]/li/text()
Bonus: Capture full text content (even with mixed elements)
If you want to ensure you get all text inside the <li> (including text nested in child elements), use the string() function:string(//h4[normalize-space(text())="Source"]/following-sibling::ul[1]/li)
This returns the concatenated text content of the node, regardless of any nested tags.
内容的提问来源于stack exchange,提问作者Novienta

