You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用XSLT 1.0将扁平XML转换为层级结构XML?

用XSLT 1.0实现扁平XML到层级XML的转换

这确实是个有点挑战性的转换需求——把平级排列的章节节点(sec1/sec2/sec3)转换成嵌套的层级结构,还要生成带层级的格式化ID。在XSLT 1.0里,我们可以利用节点轴(preceding-sibling/following-sibling)判断章节的归属关系,结合递归模板和格式化函数来实现。

输入的扁平XML

<document>
  <sec1>heading (depth 1)</sec1>
  <p>body</p>
  <sec1>heading (depth 1)</sec1>
  <sec2>heading (depth 2)</sec2>
  <p>body</p>
  <sec1>heading (depth 1)</sec1>
  <sec2>heading (depth 2)</sec2>
  <sec3>heading (depth 3)</sec3>
  <p>body</p>
</document>

实现的XSLT 1.0代码

<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="xml" indent="yes" encoding="UTF-8"/>
  
  <!-- 匹配根节点document,处理所有顶级sec1 -->
  <xsl:template match="document">
    <document>
      <xsl:apply-templates select="sec1"/>
    </document>
  </xsl:template>
  
  <!-- 处理sec1节点 -->
  <xsl:template match="sec1">
    <!-- 生成sec1的ID:三位带前导零的序号 -->
    <sec1 id="{format-number(count(preceding-sibling::sec1) + 1, '000')}">
      <!-- 将标题首字母大写 -->
      <title>
        <xsl:value-of select="translate(substring(., 1, 1), 'abcdefghijklmnopqrstuvwxyz', 'ABCDEFGHIJKLMNOPQRSTUVWXYZ')"/>
        <xsl:value-of select="substring(., 2)"/>
      </title>
      <!-- 处理当前sec1所属的后续节点(直到下一个sec1之前) -->
      <xsl:apply-templates select="following-sibling::node()[not(self::sec1) and generate-id(preceding-sibling::sec1[1]) = generate-id(current())]"/>
    </sec1>
  </xsl:template>
  
  <!-- 处理sec2节点 -->
  <xsl:template match="sec2">
    <!-- 获取父sec1的ID -->
    <xsl:variable name="parent-id" select="preceding-sibling::sec1[1]/@id"/>
    <!-- 生成sec2的序号:当前sec1下的sec2数量+1 -->
    <xsl:variable name="sec2-index" select="count(preceding-sibling::sec2[generate-id(preceding-sibling::sec1[1]) = generate-id(current()/preceding-sibling::sec1[1])]) + 1"/>
    <sec2 id="{$parent-id}-{$sec2-index}">
      <title>
        <xsl:value-of select="translate(substring(., 1, 1), 'abcdefghijklmnopqrstuvwxyz', 'ABCDEFGHIJKLMNOPQRSTUVWXYZ')"/>
        <xsl:value-of select="substring(., 2)"/>
      </title>
      <!-- 处理当前sec2所属的后续节点(直到下一个sec1或sec2之前) -->
      <xsl:apply-templates select="following-sibling::node()[not(self::sec1 or self::sec2) and generate-id(preceding-sibling::sec2[1]) = generate-id(current())]"/>
    </sec2>
  </xsl:template>
  
  <!-- 处理sec3节点 -->
  <xsl:template match="sec3">
    <!-- 获取父sec2的ID -->
    <xsl:variable name="parent-id" select="preceding-sibling::sec2[1]/@id"/>
    <!-- 生成sec3的序号:当前sec2下的sec3数量+1 -->
    <xsl:variable name="sec3-index" select="count(preceding-sibling::sec3[generate-id(preceding-sibling::sec2[1]) = generate-id(current()/preceding-sibling::sec2[1])]) + 1"/>
    <sec3 id="{$parent-id}-{$sec3-index}">
      <title>
        <xsl:value-of select="translate(substring(., 1, 1), 'abcdefghijklmnopqrstuvwxyz', 'ABCDEFGHIJKLMNOPQRSTUVWXYZ')"/>
        <xsl:value-of select="substring(., 2)"/>
      </title>
      <!-- 处理当前sec3所属的后续节点(直到下一个sec1/sec2/sec3之前) -->
      <xsl:apply-templates select="following-sibling::node()[not(self::sec1 or self::sec2 or self::sec3) and generate-id(preceding-sibling::sec3[1]) = generate-id(current())]"/>
    </sec3>
  </xsl:template>
  
  <!-- 直接复制p节点 -->
  <xsl:template match="p">
    <xsl:copy-of select="."/>
  </xsl:template>
  
  <!-- 忽略空白文本节点 -->
  <xsl:template match="text()[normalize-space()='']"/>
</xsl:stylesheet>

转换后的输出XML

<document>
  <sec1 id="001">
    <title>Heading (depth 1)</title>
    <p>body</p>
  </sec1>
  <sec1 id="002">
    <title>Heading (depth 1)</title>
    <sec2 id="002-1">
      <title>Heading (depth 2)</title>
      <p>body</p>
    </sec2>
  </sec1>
  <sec1 id="003">
    <title>Heading (depth 1)</title>
    <sec2 id="003-1">
      <title>Heading (depth 2)</title>
      <sec3 id="003-1-1">
        <title>Heading (depth 3)</title>
        <p>body</p>
      </sec3>
    </sec2>
  </sec1>
</document>

关键逻辑解释

  1. 层级归属判断:通过preceding-sibling::sec1[1]定位每个子章节(sec2/sec3)的父章节,用generate-id()确保节点匹配的唯一性,避免跨层级错误关联。
  2. ID生成:
    • sec1的ID是全局序号,用format-number()补零成三位格式(如001)。
    • 子章节ID是父章节ID加上-加上当前父层级下的局部序号(如002-1)。
  3. 递归处理:每个章节模板自动处理其所属的后续节点,嵌套子章节和内容节点,形成层级结构。
  4. 标题格式化:用translate()把标题首字母转为大写,匹配你想要的输出格式。

内容的提问来源于stack exchange,提问作者Yong Han Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:26:20