You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XSLT 2.0将XML扁平编号索引转换为树形结构?

XML扁平编号索引转树形结构 - XSLT

我需要将带编号的章节内容转换为逻辑嵌套的索引树结构,使用XSLT 2.0实现。尝试过for-each-group...group-starting-with方法,但效果不理想。

示例输入XML

<DOC>
  <section class="AB">
    <h1>Real section header</h1>
    <p><b>1. heading</b></p>
    <p>Some text here.</p>
    <p>More text.</p>
    <p><b>1.1. setting</b></p>
    <p>More words.</p>
    <p><b>1.2. fremmer</b></p>
    <p><b>1.2.1. point</b></p>
    <p>We are sailing.</p>
    <p>Whisky in the jar.</p>
    <p><b>1.2.2.</b></p>
    <p>Johnny is the man.</p>
    <p><b>1.2.3.</b></p>
    <p>And we go on and on.</p>
    <ul>
      <li>List item one</li>
      <li>List item two</li>
      <li>List item three</li>
    </ul>
    <p><b>2. Another heading</b></p>
    <p>Here is the accompanying text.</p>
    <table>
      <tr>
        <td>1</td>
        <td>Bla bla bla.</td>
      </tr>
      <tr>
        <td>2</td>
        <td>BlaX bla bla.</td>
      </tr>
      <tr>
        <td>3</td>
        <td>BlaY bla bla.</td>
      </tr>
    </table>
    <p><b>3. Last heading</b></p>
    <p>Here is the accompanying text right now.</p>
  </section>
</DOC>

期望输出XML

<DOC>
  <section class="AB">
    <h1>Real section header</h1>
    <section>
      <h1>1. heading</h1>
      <p>Some text here.</p>
      <p>More text.</p>
      <section>
        <h1>1.1. setting</h1>
        <p>More words.</p>
      </section>
      <section>
        <h1>1.2. fremmer</h1>
        <section>
          <h1>1.2.1. underpunkt</h1>
          <p>We are sailling.</p>
          <p>Whisky in the jar.</p>
        </section>
        <section>
          <h1>1.2.2.</h1>
          <p>Johnny is the man.</p>
        </section>
        <section>
          <h1>1.2.3.</h1>
          <p>And we go on and on.</p>
          <ul>
            <li>List item one</li>
            <li>List item two</li>
            <li>List item three</li>
          </ul>
        </section>
      </section>
    </section>
    <section>
      <h1>2. Another heading</h1>
      <p>Here is the accompanying text.</p>
      <table>
        <tr>
          <td>1</td>
          <td>Bla bla bla.</td>
        </tr>
        <tr>
          <td>2</td>
          <td>BlaX bla bla.</td>
        </tr>
        <tr>
          <td>3</td>
          <td>BlaY bla bla.</td>
        </tr>
      </table>
    </section>
    <section>
      <h1>3. Last heading</h1>
      <p>Here is the accompanying text right now.</p>
    </section>
  </section>
</DOC>

解决方案

核心思路是通过递归模板处理标题层级:提取标题编号判断层级,将标题到下一个同级/更高级标题间的内容作为当前section的子节点,同时递归处理子层级标题。

以下是实现的XSLT代码:

<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="xml" indent="yes"/>

  <!-- 复制根节点 -->
  <xsl:template match="/">
    <xsl:copy>
      <xsl:apply-templates/>
    </xsl:copy>
  </xsl:template>

  <!-- 处理顶层section,保留属性和h1,调用递归模板处理章节 -->
  <xsl:template match="section[@class='AB']">
    <xsl:copy>
      <xsl:copy-of select="@*"/>
      <xsl:copy-of select="h1"/>
      <xsl:call-template name="process-sections">
        <xsl:with-param name="nodes" select="node()[not(self::h1)]"/>
        <xsl:with-param name="current-level" select="1"/>
      </xsl:call-template>
    </xsl:copy>
  </xsl:template>

  <!-- 递归处理章节的核心模板 -->
  <xsl:template name="process-sections">
    <xsl:param name="nodes"/>
    <xsl:param name="current-level"/>

    <xsl:if test="$nodes">
      <!-- 定位当前层级的第一个标题节点 -->
      <xsl:variable name="current-heading" select="$nodes[1][self::p/b[matches(., '^(\d+\.)+\s?.*$')]]"/>
      
      <xsl:if test="$current-heading">
        <!-- 提取标题文本和层级 -->
        <xsl:variable name="heading-text" select="$current-heading/b"/>
        <xsl:variable name="heading-level" select="count(tokenize(normalize-space($heading-text), '\.')) - 1"/>
        
        <!-- 收集当前标题到下一个同级/更高级标题之间的内容 -->
        <xsl:variable name="content" select="$current-heading/following-sibling::node()[
          not(self::p/b[
            count(tokenize(normalize-space(.), '\.')) - 1 &lt; $heading-level 
            or (count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level and . &gt; $heading-text)
          ])
        ]"/>
        
        <section>
          <h1><xsl:value-of select="$heading-text"/></h1>
          <!-- 复制当前标题下的非标题内容 -->
          <xsl:copy-of select="$content[not(self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 &gt;= $heading-level])]"/>
          <!-- 递归处理子层级标题 -->
          <xsl:call-template name="process-sections">
            <xsl:with-param name="nodes" select="$content[self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level + 1]]"/>
            <xsl:with-param name="current-level" select="$heading-level + 1"/>
          </xsl:call-template>
        </section>
        
        <!-- 处理下一个同级标题 -->
        <xsl:call-template name="process-sections">
          <xsl:with-param name="nodes" select="$current-heading/following-sibling::node()[
            self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level and . &gt; $heading-text]
          ]"/>
          <xsl:with-param name="current-level" select="$heading-level"/>
        </xsl:call-template>
      </xsl:if>
    </xsl:if>
  </xsl:template>

  <!-- 默认复制模板,处理未匹配的节点 -->
  <xsl:template match="@*|node()">
    <xsl:copy>
      <xsl:apply-templates select="@*|node()"/>
    </xsl:copy>
  </xsl:template>
</xsl:stylesheet>

代码说明

  1. 模板匹配:保留顶层section的属性和h1标题,触发递归处理章节内容。
  2. 层级判断:通过tokenize拆分编号,统计.的数量计算标题层级,区分同级、子级和父级标题。
  3. 内容收集:筛选当前标题到下一个同级/更高级标题之间的所有内容,作为当前section的子节点。
  4. 递归处理:对当前标题下的子层级标题重复执行嵌套逻辑,生成树形结构。

内容的提问来源于stack exchange,提问作者D_Read

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 11:57:09