XSLT 1.0 如何按指定节点调整cac:InvoiceLine行排列顺序
UBL发票行按Item类型自定义排序实现
需求是基于cac:AdditionalItemProperty/cbc:Value节点存储的Item类型值,调整cac:InvoiceLine节点的排列顺序:所有类型为CU的行排在最前,类型为RC的行统一排在末尾,其余行保留原有相对顺序放在两组中间。
核心排序逻辑
给不同类型的行分配排序权重,权重值越小排列位置越靠前:
- CU类型行:权重0
- 非CU/RC的普通类型行:权重1
- RC类型行:权重2
不管用什么技术栈处理XML,只要按这个权重给行分组后按顺序拼接,就能得到符合要求的结果。
方案1:XSLT转换(最适配UBL场景的通用方案)
UBL格式的发票转换优先用XSLT,不需要额外依赖,直接保留原文档所有结构、属性和命名空间,核心实现代码如下:
<!-- 声明UBL常用命名空间,根据你实际使用的UBL版本调整对应的命名空间地址即可 --> <xsl:stylesheet version="2.0" xmlns:cac="urn:oasis:names:specification:ubl:schema:xsd:CommonAggregateComponents-2" xmlns:cbc="urn:oasis:names:specification:ubl:schema:xsd:CommonBasicComponents-2" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <!-- 恒等转换模板:原样复制所有未单独匹配的节点、属性 --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- 匹配发票根节点,单独处理InvoiceLine的排序逻辑 --> <xsl:template match="cac:Invoice"> <xsl:copy> <!-- 先复制根节点所有属性、以及非InvoiceLine的子节点,完全保留原有文档结构 --> <xsl:apply-templates select="@*|node()[not(self::cac:InvoiceLine)]"/> <!-- 遍历所有InvoiceLine节点,按自定义权重排序 --> <xsl:for-each select="cac:InvoiceLine"> <xsl:variable name="currentItemType" select="normalize-space(cac:Item/cac:AdditionalItemProperty[cbc:Name='ItemType']/cbc:Value)"/> <xsl:sort select="if ($currentItemType = 'CU') then 0 else if ($currentItemType = 'RC') then 2 else 1" data-type="number"/> <xsl:apply-templates select="."/> </xsl:for-each> </xsl:copy> </xsl:template> </xsl:stylesheet>
方案2:Python脚本快速处理
如果需要嵌入到现有Python流程里处理,可以直接用标准库xml.etree.ElementTree实现,逻辑非常直接:
import xml.etree.ElementTree as ET # 提前注册UBL命名空间,避免输出时自动生成ns0这类临时前缀 UBL_NS = { "cac": "urn:oasis:names:specification:ubl:schema:xsd:CommonAggregateComponents-2", "cbc": "urn:oasis:names:specification:ubl:schema:xsd:CommonBasicComponents-2" } for prefix, uri in UBL_NS.items(): ET.register_namespace(prefix, uri) # 读取原发票文件 tree = ET.parse("your_input_invoice.xml") invoice_root = tree.getroot() # 取出所有发票行,按类型分三组 all_lines = invoice_root.findall("cac:InvoiceLine", UBL_NS) cu_group = [] normal_group = [] rc_group = [] for line in all_lines: # 提取当前行的Item类型值,根据实际XML结构调整XPath路径 type_value = line.findtext("cac:Item/cac:AdditionalItemProperty/cbc:Value", default="", namespaces=UBL_NS).strip() if type_value == "CU": cu_group.append(line) elif type_value == "RC": rc_group.append(line) else: normal_group.append(line) # 删除原有发票行节点 for line in all_lines: invoice_root.remove(line) # 按CU组 -> 普通组 -> RC组的顺序重新插入节点 for sorted_line in cu_group + normal_group + rc_group: invoice_root.append(sorted_line) # 输出排序后的文件 tree.write("sorted_invoice.xml", encoding="utf-8", xml_declaration=True)
注意:如果你的XML里cac:AdditionalItemProperty存在多个节点,需要通过兄弟节点cbc:Name的取值筛选出存储Item类型的那个节点再取cbc:Value,避免取错值导致排序异常。
内容的提问来源于stack exchange,提问作者Domagoj Salopek
相关产品推荐
相关产品推荐

