如何用XPath获取包含子元素的完整<desc>元素及全部<desc>节点?
解决XPath获取元素完整文本(含子元素内容)的问题
问题说明
需要获取xml:id为sooke的<place>节点下所有<desc>元素的完整文本内容,包括<name>、<bibl>等子元素内的文本。当前使用的XPath语句:
if //*[@xml:id='sooke'] then ('Description:', //*[@xml:id='sooke']/desc) else ('the place does not exist')
仅能匹配到<desc>节点,但无法提取包含子元素的完整文本,尝试descendant-or-self::变体也未解决。
对应的XML示例:
<place xml:id="sooke"> <placeName>Sooke</placeName> <location> <geo>48.377315 -123.723832</geo> </location> <desc>In 1842, <name key="douglas_j">James Douglas</name> refers to Sooke as "Sy-yousung," and makes several entries about the geographical features of Sooke in <ref type="doc" cRef="V465HB02.scx">this despatch</ref>. In another spelling, with the addition of the letter "i," <name key="douglas_j">Douglas</name> refers to "Sy-yousuing" again in an <ref type="doc" cRef="V515HB02.scx">1849 despatch</ref>, wherein he states that it is "25 miles distant from <name type="place" key="victoria" >Fort Victoria</name>," and "has the important advantage of a good mill stream and a great abundance of fine timber."</desc> <desc>Another possible Sooke-landscape reference exists in the name "<ref type="external" target="http://bcgenesis.uvic.ca/getDoc.htm?id=V465HB02.scx&search=whoyring#searchHit1" >Whoyring</ref>," present day <name type="place">Becher Bay</name>, which <name key="douglas_j">Douglas</name> refers to as a port, located eight miles east of "Sy-yousuing" (<bibl>G.P.V. Akrigg and H.B. Akrigg, <title level="m">British Columbia Chronicle, 1788-1846</title> (Victoria: Discovery Press, 1975), 349</bibl>).</desc> <!-- <name type="place" key="sooke">Sooke</name> --> </place>
核心原因
你的XPath返回的是<desc>节点对象本身,而非节点包含的所有文本内容(包括子节点文本)。要提取完整文本,必须明确指定获取节点的字符串值。
解决方案
1. XPath 1.0(多数工具默认版本)
XPath 1.0中,string()函数会自动拼接单个节点下所有后代的文本内容。针对两个<desc>节点,可按以下方式处理:
- 获取第一个
<desc>的完整文本:
string(//*[@xml:id='sooke']/desc[1])
- 获取第二个
<desc>的完整文本:
string(//*[@xml:id='sooke']/desc[2])
- 结合条件判断的完整语句(用换行符分隔两个描述):
if (//*[@xml:id='sooke']) then concat('Description 1: ', string(//*[@xml:id='sooke']/desc[1]), ' Description 2: ', string(//*[@xml:id='sooke']/desc[2])) else 'the place does not exist'
注: 是XML转义后的换行符,可根据使用工具的支持情况调整。
2. XPath 2.0+(支持节点集迭代)
如果你的工具支持XPath 2.0或更高版本,可以更简洁地批量处理所有<desc>节点:
- 获取所有
<desc>文本并换行分隔:
if (//*[@xml:id='sooke']) then string-join(//*[@xml:id='sooke']/desc/string(), ' ') else 'the place does not exist'
- 带序号前缀的版本:
if (//*[@xml:id='sooke']) then string-join(for $d in //*[@xml:id='sooke']/desc return concat('Description ', position(), ': ', string($d)), ' ') else 'the place does not exist'
关于descendant-or-self::的说明
你之前尝试的descendant-or-self::如果仅用于匹配节点,仍需配合string()提取文本。比如string(//*[@xml:id='sooke']/desc/descendant-or-self::text())和直接string(//*[@xml:id='sooke']/desc)效果完全一致,因为string()默认会提取节点下所有后代文本。
内容的提问来源于stack exchange,提问作者ghosty
相关产品推荐
相关产品推荐

