You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XSLT 2/3嵌套分组:用for-each-group保持输入顺序的实现需求

XSLT列表嵌套处理:保持输入顺序的分组解决方案

问题描述

Stack Overflow上有类似问题,但现有答案会对匹配项排序或分组,导致输出顺序与输入不符。需仅通过XSLT处理输入XML,实现特定规则的列表嵌套,无法修改输入数据。

输入数据

<doc>
<paragraph indent="0" stylename="heading_l1_toc" SDgroup="heading">heading-l1</paragraph>
<paragraph indent="0" stylename="heading_l2_toc" SDgroup="heading">heading-l2</paragraph>
<paragraph indent="0" stylename="list" SDgroup="list">AAA. Para 1.</paragraph>
<paragraph indent="1" stylename="list" SDgroup="list">BBB. Para 2</paragraph>
<paragraph indent="2" stylename="continued" SDgroup="list">BBB. Continued para1</paragraph>
<paragraph indent="2" stylename="continued" SDgroup="list">BBB. Continued para2</paragraph>
<paragraph indent="2" stylename="list" SDgroup="list">CCC. Para 3a</paragraph>
<paragraph indent="2" stylename="list" SDgroup="list">DDD. Para 3b/4</paragraph>
<paragraph indent="1" stylename="list" SDgroup="list">EEE. Para 5</paragraph>
<paragraph indent="2" stylename="continued" SDgroup="list">EEE. Continued para</paragraph>
<paragraph indent="0" stylename="list" SDgroup="list">FFF. Para 6</paragraph>
</doc>

需求规则

基于list/@indent = x和continued/@indent = x + 1的规则,对指定段落进行嵌套处理:

  • 将所有stylename="list"的段落,stylename替换为list_unordered
  • 按缩进层级嵌套生成<List>节点,每个<List>需包含对应层级的indent属性,以及统计直接子项数量的maxItems属性
  • 严格保留输入的原始顺序,不得打乱或重复数据

预期输出

<doc>
  <paragraph indent="0" stylename="heading_l1_toc" SDgroup="heading">heading-l1</paragraph>
  <paragraph indent="0" stylename="heading_l2_toc" SDgroup="heading">heading-l2</paragraph>

  <List indent="0" maxItems="2">
    <paragraph stylename="list_unordered" SDgroup="list">AAA. Para 1.</paragraph>
      <List indent="1" maxItems="1">
        <paragraph indent="1" stylename="list_unordered" SDgroup="list">BBB. Para 2</paragraph>
        <paragraph indent="2" stylename="continued" SDgroup="list">BBB. Continued para1</paragraph>
        <paragraph indent="2" stylename="continued" SDgroup="list">BBB. Continued para2</paragraph>
        <List indent="2" maxItems="2">
          <paragraph indent="2" stylename="list_unordered" SDgroup="list">CCC. Para 3a</paragraph>
          <paragraph indent="2" stylename="list_unordered" SDgroup="list">DDD. Para 34</paragraph>
        </List>
        <paragraph indent="1" stylename="list_unordered" SDgroup="list">EEE. Para 5</paragraph>
        <paragraph indent="2" stylename="continued" SDgroup="list">EEE. Continued para</paragraph>
      </List>
      <paragraph indent="0" stylename="list_unordered" SDgroup="list">FFF. Para 6</paragraph>
  </List>
</doc>

已尝试方案

试过多种for-each-group结合group-by/adjacent以及普通for-each的写法,但要么输出顺序混乱,要么嵌套正确但出现数据重复。


解决方案

采用递归处理逻辑,按输入顺序逐个处理节点,根据缩进层级判断嵌套时机,同时统计列表项数量,完美保留原始顺序:

<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
    <xsl:output method="xml" indent="yes"/>

    <!-- 处理根节点,先输出标题,再启动列表递归处理 -->
    <xsl:template match="/doc">
        <doc>
            <xsl:copy-of select="paragraph[@SDgroup='heading']"/>
            <xsl:text>&#10;&#10;</xsl:text>
            <xsl:call-template name="process-list">
                <xsl:with-param name="current-nodes" select="paragraph[@stylename='list' or @stylename='continued']"/>
                <xsl:with-param name="current-indent" select="0"/>
            </xsl:call-template>
        </doc>
    </xsl:template>

    <!-- 递归处理列表的核心模板 -->
    <xsl:template name="process-list">
        <xsl:param name="current-nodes"/>
        <xsl:param name="current-indent"/>
        
        <!-- 筛选当前层级的list节点作为直接子项 -->
        <xsl:variable name="list-items" select="$current-nodes[@stylename='list' and @indent = $current-indent]"/>
        
        <xsl:if test="$list-items">
            <List indent="{$current-indent}" maxItems="{count($list-items)}">
                <xsl:for-each select="$list-items">
                    <!-- 输出替换stylename后的list节点 -->
                    <paragraph>
                        <xsl:if test="@indent">
                            <xsl:attribute name="indent" select="@indent"/>
                        </xsl:if>
                        <xsl:attribute name="stylename" select="'list_unordered'"/>
                        <xsl:copy-of select="@SDgroup"/>
                        <xsl:value-of select="."/>
                    </paragraph>
                    
                    <!-- 定位当前list后续的子层级节点 -->
                    <xsl:variable name="following-nodes" select="following-sibling::paragraph[
                        (@stylename='continued' and @indent = $current-indent + 1) or
                        (@stylename='list' and @indent > $current-indent)
                    ][
                        not(following-sibling::paragraph[@stylename='list' and @indent <= $current-indent][1] << .)
                    ]"/>
                    
                    <!-- 递归处理子层级 -->
                    <xsl:if test="$following-nodes">
                        <xsl:call-template name="process-list">
                            <xsl:with-param name="current-nodes" select="$following-nodes"/>
                            <xsl:with-param name="current-indent" select="$current-indent + 1"/>
                        </xsl:call-template>
                    </xsl:if>
                </xsl:for-each>
            </List>
        </xsl:if>
    </xsl:template>

    <!-- 匹配无需处理的节点,默认不输出 -->
    <xsl:template match="paragraph"/>
</xsl:stylesheet>

方案说明

  1. 递归逻辑:从最外层(indent=0)开始,逐层处理子层级,完全遵循输入节点的原始顺序
  2. 精准筛选:通过following-sibling和位置判断,准确匹配当前list对应的后续子节点(包括continued和更深层级的list)
  3. 属性处理:自动替换list节点的stylename,保留其他必要属性
  4. 数量统计:直接统计当前层级的list节点数量,生成maxItems属性
  5. 顺序保障:全程无排序或分组操作,严格按照输入顺序输出内容

内容的提问来源于stack exchange,提问作者markh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 16:44:51