如何用XSLT在大型XML文档中包裹指定相邻节点序列?
解决方案:用XSLT匹配相邻节点序列并包裹
完整XSLT代码
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <!-- 身份模板:复制所有未被特殊处理的节点 --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- 匹配符合条件的<a>元素,将目标序列包裹到<bar>中 --> <xsl:template match="a[ following-sibling::node()[1][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[2][self::b] and following-sibling::node()[3][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[4][self::c] ]"> <bar> <!-- 复制当前<a>及后续的目标节点 --> <xsl:copy-of select="."/> <xsl:copy-of select="following-sibling::node()[1]"/> <xsl:copy-of select="following-sibling::node()[2]"/> <xsl:copy-of select="following-sibling::node()[3]"/> <xsl:copy-of select="following-sibling::node()[4]"/> </bar> </xsl:template> <!-- 跳过已被包裹到<bar>中的节点,避免重复输出 --> <xsl:template match="text()[ preceding-sibling::a[1][ following-sibling::node()[1][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[2][self::b] and following-sibling::node()[3][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[4][self::c] ] ]"/> <xsl:template match="b[ preceding-sibling::a[1][ following-sibling::node()[1][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[2][self::b] and following-sibling::node()[3][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[4][self::c] ] ]"/> <xsl:template match="c[ preceding-sibling::a[2][ following-sibling::node()[1][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[2][self::b] and following-sibling::node()[3][self::text()[contains(normalize-space(), ',')]] and following-sibling::node()[4][self::c] ] ]"/> </xsl:stylesheet>
关键说明
- 身份模板:作为基础规则,原样复制XML中所有未被特殊模板匹配的节点,保证文档其他内容不受改动。
- 目标序列匹配:第一个模板精准定位符合要求的
<a>元素——它的后续兄弟节点必须依次是带逗号的文本节点、<b>元素、带逗号的文本节点、<c>元素。匹配成功后,将这一组节点全部包裹到<bar>标签内。 - 避免重复输出:后续三个模板会跳过已被纳入
<bar>的节点(对应序列里的文本节点、<b>、<c>),防止这些节点被身份模板再次复制输出。 - 文本兼容处理:使用
normalize-space()配合contains(),可以兼容文本节点中逗号前后的空格差异(比如, '或者, '等格式)。如果你的文本格式完全固定,也可以直接用contains(., ', ')实现更精准的匹配。
内容的提问来源于stack exchange,提问作者cyocum
相关产品推荐
相关产品推荐

