You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XPath获取包含子元素的完整<desc>元素及全部<desc>节点?

解决XPath获取元素完整文本(含子元素内容)的问题

问题说明

需要获取xml:id为sooke的<place>节点下所有<desc>元素的完整文本内容,包括<name>、<bibl>等子元素内的文本。当前使用的XPath语句:

if //*[@xml:id='sooke'] then ('Description:', //*[@xml:id='sooke']/desc) else ('the place does not exist')

仅能匹配到<desc>节点,但无法提取包含子元素的完整文本,尝试descendant-or-self::变体也未解决。

对应的XML示例:

<place xml:id="sooke">
  <placeName>Sooke</placeName>
  <location>
    <geo>48.377315 -123.723832</geo>
  </location>
  <desc>In 1842, <name key="douglas_j">James Douglas</name> refers to Sooke as
    "Sy-yousung," and makes several entries about the geographical features of Sooke in
      <ref type="doc" cRef="V465HB02.scx">this despatch</ref>. In another spelling, with
    the addition of the letter "i," <name key="douglas_j">Douglas</name> refers to
    "Sy-yousuing" again in an <ref type="doc" cRef="V515HB02.scx">1849 despatch</ref>,
    wherein he states that it is "25 miles distant from <name type="place" key="victoria"
      >Fort Victoria</name>," and "has the important advantage of a good mill stream and a
    great abundance of fine timber."</desc>
  <desc>Another possible Sooke-landscape reference exists in the name "<ref
      type="external"
      target="http://bcgenesis.uvic.ca/getDoc.htm?id=V465HB02.scx&amp;search=whoyring#searchHit1"
      >Whoyring</ref>," present day <name type="place">Becher Bay</name>, which <name
      key="douglas_j">Douglas</name> refers to as a port, located eight miles east of
    "Sy-yousuing" (<bibl>G.P.V. Akrigg and H.B. Akrigg, <title level="m">British Columbia
        Chronicle, 1788-1846</title> (Victoria: Discovery Press, 1975),
    349</bibl>).</desc>
  <!-- <name type="place" key="sooke">Sooke</name> -->
</place>

核心原因

你的XPath返回的是<desc>节点对象本身,而非节点包含的所有文本内容(包括子节点文本)。要提取完整文本,必须明确指定获取节点的字符串值。

解决方案

1. XPath 1.0(多数工具默认版本)

XPath 1.0中,string()函数会自动拼接单个节点下所有后代的文本内容。针对两个<desc>节点,可按以下方式处理:

  • 获取第一个<desc>的完整文本:
string(//*[@xml:id='sooke']/desc[1])
  • 获取第二个<desc>的完整文本:
string(//*[@xml:id='sooke']/desc[2])
  • 结合条件判断的完整语句(用换行符分隔两个描述):
if (//*[@xml:id='sooke']) then concat('Description 1: ', string(//*[@xml:id='sooke']/desc[1]), '&#10;Description 2: ', string(//*[@xml:id='sooke']/desc[2])) else 'the place does not exist'

注:&#10;是XML转义后的换行符,可根据使用工具的支持情况调整。

2. XPath 2.0+(支持节点集迭代)

如果你的工具支持XPath 2.0或更高版本,可以更简洁地批量处理所有<desc>节点:

  • 获取所有<desc>文本并换行分隔:
if (//*[@xml:id='sooke']) then string-join(//*[@xml:id='sooke']/desc/string(), '&#10;') else 'the place does not exist'
  • 带序号前缀的版本:
if (//*[@xml:id='sooke']) then string-join(for $d in //*[@xml:id='sooke']/desc return concat('Description ', position(), ': ', string($d)), '&#10;') else 'the place does not exist'

关于descendant-or-self::的说明

你之前尝试的descendant-or-self::如果仅用于匹配节点,仍需配合string()提取文本。比如string(//*[@xml:id='sooke']/desc/descendant-or-self::text())和直接string(//*[@xml:id='sooke']/desc)效果完全一致,因为string()默认会提取节点下所有后代文本。

内容的提问来源于stack exchange,提问作者ghosty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 18:12:12