You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写XPath获取XML中指定h4标题后的列表项内容?

Correct XPath Expressions for Your Case

Your initial XPath didn't work because the <ul> is a sibling of the <h4>, not a child. The <h4> and <ul> are both nested inside the same parent <div>, so we need to traverse to the following sibling <ul> instead of looking for a direct child.

Here are the precise, position-agnostic XPath expressions for each heading:

  • Source: //h4[normalize-space(text())="Source"]/following-sibling::ul[1]/li/text()

    • normalize-space(text()) handles accidental whitespace in the h4 text (like leading/trailing spaces)
    • following-sibling::ul[1] targets the very first <ul> immediately after the matching <h4>, ensuring we grab the correct associated list even if other elements are nearby
  • Language: //h4[normalize-space(text())="Language"]/following-sibling::ul[1]/li/text()

    • This captures the text inside the <a> tag (e.g., "Swedish" in your example) because text() retrieves all text nodes within the <li>
  • Publishers: //h4[normalize-space(text())="Publishers"]/following-sibling::ul[1]/li/text()

  • Record ID: //h4[normalize-space(text())="Record ID"]/following-sibling::ul[1]/li/text()

Bonus: Capture full text content (even with mixed elements)

If you want to ensure you get all text inside the <li> (including text nested in child elements), use the string() function:
string(//h4[normalize-space(text())="Source"]/following-sibling::ul[1]/li)
This returns the concatenated text content of the node, regardless of any nested tags.

内容的提问来源于stack exchange,提问作者Novienta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:21:01