You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用XSLT for-each-group实现XML节点分组封装的方案优化

使用XSLT for-each-group实现step节点的容器分组需求

先明确你的核心需求:识别仅包含<w>元素的<step>节点,将其转换为<container>,并把后续相邻的不含<w>的<step>节点全部纳入这个容器;而那些同时包含<w>和其他元素(比如<p>)的<step>则保持独立,不参与分组。

需求规则回顾

  • 匹配仅包含<w>元素的<step>节点,将其转为<container>,并收纳后续相邻的无<w>的<step>
  • 含<w>+其他元素的<step>节点,直接保留原样
  • 无<w>的<step>节点:如果前面是仅含<w>的<step>,则被纳入容器;否则直接保留

输入与期望输出示例

输入XML片段

<?xml version="1.0" encoding="UTF-8"?>
<body>
<step>
<p>step 1</p>
</step>
<step>
<p>step 2</p>
</step>
<step>
<w>Warning A</w>
<p>step 3</p>
</step>
<step>
<p>step 4</p>
</step>
<step>
<p>step 5</p>
</step>
<step>
<w>Spec Warning X</w>
<w>Spec Warning Y</w>
</step>
<step>
<p>step 6</p>
</step>
<step>
<p>step 7</p>
</step>
<step>
<p>step 8</p>
</step>
<step>
<p>step 9</p>
</step>
<step>
<p>step 10</p>
</step>
<step>
<p>step 11</p>
</step>
<step>
<w>Warning B</w>
<p>step 12</p>
</step>
<step>
<p>step 13</p>
</step>
<step>
<p>step 14</p>
</step>
</body>

期望输出XML片段

<?xml version="1.0" encoding="UTF-8"?>
<body>
<step>
<p>step 1</p>
</step>
<step>
<p>step 2</p>
</step>
<step>
<w>Warning A</w>
<p>step 3</p>
</step>
<step>
<p>step 4</p>
</step>
<step>
<p>step 5</p>
</step>
<container>
<w>Spec Warning X</w>
<w>Spec Warning Y</w>
<step>
<p>step 6</p>
</step>
<step>
<p>step 7</p>
</step>
<step>
<p>step 8</p>
</step>
<step>
<p>step 9</p>
</step>
<step>
<p>step 10</p>
</step>
<step>
<p>step 11</p>
</step>
</container>
<step>
<w>Warning B</w>
<p>step 12</p>
</step>
<step>
<p>step 13</p>
</step>
<step>
<p>step 14</p>
</step>
</body>

对两次尝试的点评

第一次尝试的问题

你第一次写的代码里,group-adjacent="self::step[not(w)]"会报错XTTE1100,原因是group-adjacent要求返回一个单一值(比如布尔值、字符串),而你写的表达式返回的是节点集/布尔序列,不符合语法要求。另外,你没有限制只取当前仅含<w>的<step>之后的第一个连续组,而是把所有后续节点都分组,这会导致逻辑混乱。

第二次尝试的优缺点

第二次尝试里,你把group-adjacent改成了boolean(self::step[not(w)]),这就符合语法要求了——返回的是布尔值,用来区分"是无w的step"和"不是"的节点。然后通过preceding-sibling::step[w][1][not(p)]来筛选目标组,这个思路是对的,但有几个小问题:

  • 模板匹配的逻辑有点绕,比如step[p][not(preceding-sibling::step[w][1][not(p)])]虽然能工作,但可读性差
  • for-each-group遍历了所有后续节点,效率不高;其实我们只需要取到第一个非无w的step之前的所有节点即可
  • 存在冗余的模板(比如单独匹配<w>的模板其实没必要,因为apply-templates会默认复制节点)

基于for-each-group的正确实现思路

核心思路是对所有<step>节点按"分组触发条件"进行分组:把仅含<w>的<step>作为分组的起始标记,后续的无<w>的<step>都归到这个组里;其他节点单独成组。

具体步骤:

  1. 对<body>下的所有<step>节点做分组,分组键用if (self::step[w and not(*) except w]) then 'container-start' else 'normal'——意思是:如果当前step只有w元素,标记为容器起始,否则标记为普通节点
  2. 遍历每个分组:
    • 如果是普通节点组,直接复制所有节点
    • 如果是容器起始组:
      • 创建<container>元素,先把当前组里的仅含w的step的内容(也就是所有w元素)复制进去
      • 然后找到这个组之后的连续的普通节点组(无w的step),把这些节点也复制到容器里
      • 注意要跳过已经被纳入容器的普通节点,避免重复输出

完整的XSLT代码

<?xml version="1.0" encoding="UTF-8"?>
<xsl:transform xmlns:xsl="http://www.w3.org/1999/XSL/Transform" version="2.0">
  <xsl:output method="xml" omit-xml-declaration="no" encoding="UTF-8" indent="yes" />

  <!-- 匹配根节点,处理body -->
  <xsl:template match="/body">
    <xsl:copy>
      <xsl:for-each-group select="step" group-adjacent="if (self::step[w and not(* except w)]) then 'container' else 'normal'">
        <xsl:choose>
          <!-- 处理普通节点组:直接复制 -->
          <xsl:when test="current-grouping-key() = 'normal'">
            <xsl:copy-of select="current-group()"/>
          </xsl:when>
          <!-- 处理容器起始组:创建container并收纳后续连续普通组 -->
          <xsl:when test="current-grouping-key() = 'container'">
            <container>
              <!-- 复制当前组中step的所有w元素 -->
              <xsl:copy-of select="current-group()/w"/>
              <!-- 找到后续第一个连续的normal组并复制 -->
              <xsl:variable name="next-normal-group" select="current-group()[last()]/following-sibling::step[not(w)]"/>
              <xsl:copy-of select="$next-normal-group"/>
            </container>
            <!-- 跳过已经被纳入容器的normal组 -->
            <xsl:apply-templates select="$next-normal-group" mode="skip"/>
          </xsl:when>
        </xsl:choose>
      </xsl:for-each-group>
    </xsl:copy>
  </xsl:template>

  <!-- 空模板:跳过已经被处理的normal节点 -->
  <xsl:template match="step" mode="skip"/>

  <!-- 默认模板:复制其他节点(比如w、p等) -->
  <xsl:template match="@*|node()">
    <xsl:copy>
      <xsl:apply-templates select="@*|node()"/>
    </xsl:copy>
  </xsl:template>
</xsl:transform>

代码解释

  • group-adjacent的分组键:self::step[w and not(* except w)]用来精准判断"仅包含w元素的step"——* except w会排除w元素,not(...)表示没有其他子元素
  • 当遇到容器起始组时,先复制该step里的所有w元素到<container>,然后通过current-group()[last()]/following-sibling::step[not(w)]获取后续所有连续的无w的step节点
  • 用mode="skip"的空模板来跳过这些已经被纳入容器的节点,避免重复输出
  • 默认模板(identity template)负责复制所有其他节点,不需要单独写模板匹配<w>或<p>

这样的实现逻辑清晰,效率更高,完全符合你的需求,而且是for-each-group的标准用法。

内容的提问来源于stack exchange,提问作者Zug_Bug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:29:11