使用XSLT转换XML分组节点保留注释且无需预知其余标签名的方法
XML重排需求:仅已知目标节点标签时用XSLT实现分组排序
我有一个结构简单的XML文件,希望通过xmlstarlet对其进行重排。
示例输入
<myXml description="example 1"> <!-- Comment XXX --> <randomNodeX> <randomSubNode1>value1</randomSubNode1> <randomSubNode2>value2</randomSubNode2> </randomNodeX> <!-- Comment YYY1 --> <!-- Comment YYY2 --> <randomNodeY attribute1="value3" attribute2="value4"/> <!-- Comment ZZZ --> <randomNodeZ attribute1="value5" attribute0="value6"> <randomSubNode3 attribute3="value7" attribute4="value8"/> </randomNodeZ> <!-- Comment for node1 first occurrence --> <node1 attribute1="value9" attribute5="value10" attribute6="value11"/> <!-- Comment for node2 first occurrence --> <node2 attribute1="value12" attribute7="value13" attribute8="value14"> <subNode21 attributeX="value15"/> <subNode22 attributeY="value16" attributeZ="value17"/> </node2> <!-- Comment for node3 first occurrence --> <node3 attribute1="value18" attribute9="value19"> <subNode31 attributeW="value20"/> </node3> <!-- Comment for node1 second occurrence --> <node1 attribute1="value21" attribute5="value22" attribute6="value23"/> <!-- Comment for node3 second occurrence --> <node3 attribute1="value24" attribute9="value25"> <subNode31 attributeW="value26"/> </node3> <!-- Comment for node2 second occurrence --> <node2 attribute1="value27" attribute7="value28" attribute8="value29"> <subNode21 attributeX="value30"/> <subNode22 attributeY="value31" attributeZ="value32"/> </node2> </myXml>
重排规则
- 所有
node1、node2、node3元素需要连同各自对应的注释集中放置,按node1→node2→node3的顺序分组 - 除此之外的其余文档内容及注释全部保留在文档开头,无需提前知晓这些非目标节点的标签名
预期输出
<myXml description="example 1"> <!-- Comment XXX --> <randomNodeX> <randomSubNode1>value1</randomSubNode1> <randomSubNode2>value2</randomSubNode2> </randomNodeX> <!-- Comment YYY1 --> <!-- Comment YYY2 --> <randomNodeY attribute1="value3" attribute2="value4"/> <!-- Comment ZZZ --> <randomNodeZ attribute1="value5" attribute0="value6"> <randomSubNode3 attribute3="value7" attribute4="value8"/> </randomNodeZ> <!-- Comment for node1 first occurrence --> <node1 attribute1="value9" attribute5="value10" attribute6="value11"/> <!-- Comment for node1 second occurrence --> <node1 attribute1="value21" attribute5="value22" attribute6="value23"/> <!-- Comment for node2 first occurrence --> <node2 attribute1="value12" attribute7="value13" attribute8="value14"> <subNode21 attributeX="value15"/> <subNode22 attributeY="value16" attributeZ="value17"/> </node2> <!-- Comment for node2 second occurrence --> <node2 attribute1="value27" attribute7="value28" attribute8="value29"> <subNode21 attributeX="value30"/> <subNode22 attributeY="value31" attributeZ="value32"/> </node2> <!-- Comment for node3 first occurrence --> <node3 attribute1="value18" attribute9="value19"> <subNode31 attributeW="value20"/> </node3> <!-- Comment for node3 second occurrence --> <node3 attribute1="value24" attribute9="value25"> <subNode31 attributeW="value26"/> </node3> </myXml>
现有实现的问题
目前写的XSLT需要提前指定所有非目标节点的标签才能运行,希望仅已知node1/node2/node3三个标签的前提下实现需求:
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output indent="yes"/> <xsl:strip-space elements="*"/> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()[not(self::node1|self::node2|self::node3|self::comment())]"/> <xsl:apply-templates select="node1"/> <xsl:apply-templates select="node2"/> <xsl:apply-templates select="node3"/> </xsl:copy> </xsl:template> <xsl:template match="randomNodeX|randomNodeY|randomNodeZ|node1|node2|node3"> <xsl:apply-templates select="preceding-sibling::comment()[generate-id(following-sibling::*[1])=generate-id(current())]"/> <xsl:copy-of select="."/> </xsl:template> </xsl:stylesheet>
解决方案
修改XSLT的匹配规则即可,无需提前枚举非目标节点:
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output indent="yes"/> <xsl:strip-space elements="*"/> <!-- 根节点匹配,控制输出顺序 --> <xsl:template match="/*"> <xsl:copy> <xsl:apply-templates select="@*"/> <!-- 第一步:输出所有非node1/node2/node3的元素,以及它们对应的前置注释 --> <xsl:apply-templates select="*[not(self::node1 or self::node2 or self::node3)]"/> <!-- 第二步:按顺序输出分组的node1、node2、node3,带对应前置注释 --> <xsl:apply-templates select="node1"/> <xsl:apply-templates select="node2"/> <xsl:apply-templates select="node3"/> </xsl:copy> </xsl:template> <!-- 通用元素匹配:输出元素本身 + 紧邻它的前置注释 --> <xsl:template match="*"> <xsl:apply-templates select="preceding-sibling::comment()[generate-id(following-sibling::*[1]) = generate-id(current())]"/> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- 注释匹配:原样输出 --> <xsl:template match="comment()"> <xsl:copy/> </xsl:template> <!-- 属性、文本等节点默认复制 --> <xsl:template match="@*|text()|processing-instruction()"> <xsl:copy/> </xsl:template> </xsl:stylesheet>
实现说明
- 根节点模板里首先筛选所有非目标节点输出,不需要知道这些节点的具体标签名,只要不是
node1/node2/node3就会被放到开头 - 通用元素匹配规则会自动处理所有元素的前置注释提取,不需要单独枚举节点名
- 所有非元素节点(属性、文本、处理指令等)都默认原样保留,兼容性更好
调用命令参考:
xmlstarlet tr rearrange.xsl input.xml > output.xml
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

