求助:使用XSLT for-each-group实现XML节点分组封装的方案优化
使用XSLT
for-each-group实现step节点的容器分组需求 先明确你的核心需求:识别仅包含<w>元素的<step>节点,将其转换为<container>,并把后续相邻的不含<w>的<step>节点全部纳入这个容器;而那些同时包含<w>和其他元素(比如<p>)的<step>则保持独立,不参与分组。
需求规则回顾
- 匹配仅包含
<w>元素的<step>节点,将其转为<container>,并收纳后续相邻的无<w>的<step> - 含
<w>+其他元素的<step>节点,直接保留原样 - 无
<w>的<step>节点:如果前面是仅含<w>的<step>,则被纳入容器;否则直接保留
输入与期望输出示例
输入XML片段
<?xml version="1.0" encoding="UTF-8"?> <body> <step> <p>step 1</p> </step> <step> <p>step 2</p> </step> <step> <w>Warning A</w> <p>step 3</p> </step> <step> <p>step 4</p> </step> <step> <p>step 5</p> </step> <step> <w>Spec Warning X</w> <w>Spec Warning Y</w> </step> <step> <p>step 6</p> </step> <step> <p>step 7</p> </step> <step> <p>step 8</p> </step> <step> <p>step 9</p> </step> <step> <p>step 10</p> </step> <step> <p>step 11</p> </step> <step> <w>Warning B</w> <p>step 12</p> </step> <step> <p>step 13</p> </step> <step> <p>step 14</p> </step> </body>
期望输出XML片段
<?xml version="1.0" encoding="UTF-8"?> <body> <step> <p>step 1</p> </step> <step> <p>step 2</p> </step> <step> <w>Warning A</w> <p>step 3</p> </step> <step> <p>step 4</p> </step> <step> <p>step 5</p> </step> <container> <w>Spec Warning X</w> <w>Spec Warning Y</w> <step> <p>step 6</p> </step> <step> <p>step 7</p> </step> <step> <p>step 8</p> </step> <step> <p>step 9</p> </step> <step> <p>step 10</p> </step> <step> <p>step 11</p> </step> </container> <step> <w>Warning B</w> <p>step 12</p> </step> <step> <p>step 13</p> </step> <step> <p>step 14</p> </step> </body>
对两次尝试的点评
第一次尝试的问题
你第一次写的代码里,group-adjacent="self::step[not(w)]"会报错XTTE1100,原因是group-adjacent要求返回一个单一值(比如布尔值、字符串),而你写的表达式返回的是节点集/布尔序列,不符合语法要求。另外,你没有限制只取当前仅含<w>的<step>之后的第一个连续组,而是把所有后续节点都分组,这会导致逻辑混乱。
第二次尝试的优缺点
第二次尝试里,你把group-adjacent改成了boolean(self::step[not(w)]),这就符合语法要求了——返回的是布尔值,用来区分"是无w的step"和"不是"的节点。然后通过preceding-sibling::step[w][1][not(p)]来筛选目标组,这个思路是对的,但有几个小问题:
- 模板匹配的逻辑有点绕,比如
step[p][not(preceding-sibling::step[w][1][not(p)])]虽然能工作,但可读性差 for-each-group遍历了所有后续节点,效率不高;其实我们只需要取到第一个非无w的step之前的所有节点即可- 存在冗余的模板(比如单独匹配
<w>的模板其实没必要,因为apply-templates会默认复制节点)
基于for-each-group的正确实现思路
核心思路是对所有<step>节点按"分组触发条件"进行分组:把仅含<w>的<step>作为分组的起始标记,后续的无<w>的<step>都归到这个组里;其他节点单独成组。
具体步骤:
- 对
<body>下的所有<step>节点做分组,分组键用if (self::step[w and not(*) except w]) then 'container-start' else 'normal'——意思是:如果当前step只有w元素,标记为容器起始,否则标记为普通节点 - 遍历每个分组:
- 如果是普通节点组,直接复制所有节点
- 如果是容器起始组:
- 创建
<container>元素,先把当前组里的仅含w的step的内容(也就是所有w元素)复制进去 - 然后找到这个组之后的连续的普通节点组(无w的step),把这些节点也复制到容器里
- 注意要跳过已经被纳入容器的普通节点,避免重复输出
- 创建
完整的XSLT代码
<?xml version="1.0" encoding="UTF-8"?> <xsl:transform xmlns:xsl="http://www.w3.org/1999/XSL/Transform" version="2.0"> <xsl:output method="xml" omit-xml-declaration="no" encoding="UTF-8" indent="yes" /> <!-- 匹配根节点,处理body --> <xsl:template match="/body"> <xsl:copy> <xsl:for-each-group select="step" group-adjacent="if (self::step[w and not(* except w)]) then 'container' else 'normal'"> <xsl:choose> <!-- 处理普通节点组:直接复制 --> <xsl:when test="current-grouping-key() = 'normal'"> <xsl:copy-of select="current-group()"/> </xsl:when> <!-- 处理容器起始组:创建container并收纳后续连续普通组 --> <xsl:when test="current-grouping-key() = 'container'"> <container> <!-- 复制当前组中step的所有w元素 --> <xsl:copy-of select="current-group()/w"/> <!-- 找到后续第一个连续的normal组并复制 --> <xsl:variable name="next-normal-group" select="current-group()[last()]/following-sibling::step[not(w)]"/> <xsl:copy-of select="$next-normal-group"/> </container> <!-- 跳过已经被纳入容器的normal组 --> <xsl:apply-templates select="$next-normal-group" mode="skip"/> </xsl:when> </xsl:choose> </xsl:for-each-group> </xsl:copy> </xsl:template> <!-- 空模板:跳过已经被处理的normal节点 --> <xsl:template match="step" mode="skip"/> <!-- 默认模板:复制其他节点(比如w、p等) --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> </xsl:transform>
代码解释
group-adjacent的分组键:self::step[w and not(* except w)]用来精准判断"仅包含w元素的step"——* except w会排除w元素,not(...)表示没有其他子元素- 当遇到容器起始组时,先复制该step里的所有w元素到
<container>,然后通过current-group()[last()]/following-sibling::step[not(w)]获取后续所有连续的无w的step节点 - 用
mode="skip"的空模板来跳过这些已经被纳入容器的节点,避免重复输出 - 默认模板(identity template)负责复制所有其他节点,不需要单独写模板匹配
<w>或<p>
这样的实现逻辑清晰,效率更高,完全符合你的需求,而且是for-each-group的标准用法。
内容的提问来源于stack exchange,提问作者Zug_Bug
相关产品推荐
相关产品推荐

