如何用XSLT 2.0将XML扁平编号索引转换为树形结构?
XML扁平编号索引转树形结构 - XSLT
我需要将带编号的章节内容转换为逻辑嵌套的索引树结构,使用XSLT 2.0实现。尝试过for-each-group...group-starting-with方法,但效果不理想。
示例输入XML
<DOC> <section class="AB"> <h1>Real section header</h1> <p><b>1. heading</b></p> <p>Some text here.</p> <p>More text.</p> <p><b>1.1. setting</b></p> <p>More words.</p> <p><b>1.2. fremmer</b></p> <p><b>1.2.1. point</b></p> <p>We are sailing.</p> <p>Whisky in the jar.</p> <p><b>1.2.2.</b></p> <p>Johnny is the man.</p> <p><b>1.2.3.</b></p> <p>And we go on and on.</p> <ul> <li>List item one</li> <li>List item two</li> <li>List item three</li> </ul> <p><b>2. Another heading</b></p> <p>Here is the accompanying text.</p> <table> <tr> <td>1</td> <td>Bla bla bla.</td> </tr> <tr> <td>2</td> <td>BlaX bla bla.</td> </tr> <tr> <td>3</td> <td>BlaY bla bla.</td> </tr> </table> <p><b>3. Last heading</b></p> <p>Here is the accompanying text right now.</p> </section> </DOC>
期望输出XML
<DOC> <section class="AB"> <h1>Real section header</h1> <section> <h1>1. heading</h1> <p>Some text here.</p> <p>More text.</p> <section> <h1>1.1. setting</h1> <p>More words.</p> </section> <section> <h1>1.2. fremmer</h1> <section> <h1>1.2.1. underpunkt</h1> <p>We are sailling.</p> <p>Whisky in the jar.</p> </section> <section> <h1>1.2.2.</h1> <p>Johnny is the man.</p> </section> <section> <h1>1.2.3.</h1> <p>And we go on and on.</p> <ul> <li>List item one</li> <li>List item two</li> <li>List item three</li> </ul> </section> </section> </section> <section> <h1>2. Another heading</h1> <p>Here is the accompanying text.</p> <table> <tr> <td>1</td> <td>Bla bla bla.</td> </tr> <tr> <td>2</td> <td>BlaX bla bla.</td> </tr> <tr> <td>3</td> <td>BlaY bla bla.</td> </tr> </table> </section> <section> <h1>3. Last heading</h1> <p>Here is the accompanying text right now.</p> </section> </section> </DOC>
解决方案
核心思路是通过递归模板处理标题层级:提取标题编号判断层级,将标题到下一个同级/更高级标题间的内容作为当前section的子节点,同时递归处理子层级标题。
以下是实现的XSLT代码:
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes"/> <!-- 复制根节点 --> <xsl:template match="/"> <xsl:copy> <xsl:apply-templates/> </xsl:copy> </xsl:template> <!-- 处理顶层section,保留属性和h1,调用递归模板处理章节 --> <xsl:template match="section[@class='AB']"> <xsl:copy> <xsl:copy-of select="@*"/> <xsl:copy-of select="h1"/> <xsl:call-template name="process-sections"> <xsl:with-param name="nodes" select="node()[not(self::h1)]"/> <xsl:with-param name="current-level" select="1"/> </xsl:call-template> </xsl:copy> </xsl:template> <!-- 递归处理章节的核心模板 --> <xsl:template name="process-sections"> <xsl:param name="nodes"/> <xsl:param name="current-level"/> <xsl:if test="$nodes"> <!-- 定位当前层级的第一个标题节点 --> <xsl:variable name="current-heading" select="$nodes[1][self::p/b[matches(., '^(\d+\.)+\s?.*$')]]"/> <xsl:if test="$current-heading"> <!-- 提取标题文本和层级 --> <xsl:variable name="heading-text" select="$current-heading/b"/> <xsl:variable name="heading-level" select="count(tokenize(normalize-space($heading-text), '\.')) - 1"/> <!-- 收集当前标题到下一个同级/更高级标题之间的内容 --> <xsl:variable name="content" select="$current-heading/following-sibling::node()[ not(self::p/b[ count(tokenize(normalize-space(.), '\.')) - 1 < $heading-level or (count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level and . > $heading-text) ]) ]"/> <section> <h1><xsl:value-of select="$heading-text"/></h1> <!-- 复制当前标题下的非标题内容 --> <xsl:copy-of select="$content[not(self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 >= $heading-level])]"/> <!-- 递归处理子层级标题 --> <xsl:call-template name="process-sections"> <xsl:with-param name="nodes" select="$content[self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level + 1]]"/> <xsl:with-param name="current-level" select="$heading-level + 1"/> </xsl:call-template> </section> <!-- 处理下一个同级标题 --> <xsl:call-template name="process-sections"> <xsl:with-param name="nodes" select="$current-heading/following-sibling::node()[ self::p/b[count(tokenize(normalize-space(.), '\.')) - 1 = $heading-level and . > $heading-text] ]"/> <xsl:with-param name="current-level" select="$heading-level"/> </xsl:call-template> </xsl:if> </xsl:if> </xsl:template> <!-- 默认复制模板,处理未匹配的节点 --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> </xsl:stylesheet>
代码说明
- 模板匹配:保留顶层section的属性和h1标题,触发递归处理章节内容。
- 层级判断:通过
tokenize拆分编号,统计.的数量计算标题层级,区分同级、子级和父级标题。 - 内容收集:筛选当前标题到下一个同级/更高级标题之间的所有内容,作为当前section的子节点。
- 递归处理:对当前标题下的子层级标题重复执行嵌套逻辑,生成树形结构。
内容的提问来源于stack exchange,提问作者D_Read
相关产品推荐
相关产品推荐

