You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python合并两个XML文件时节点嵌套异常问题求助

合并XML文件时节点嵌套问题排查与解决

问题描述

需合并两个XML文件,使两个<invoiceCopy>节点并列,但运行代码后第二个节点被嵌套在第一个内部。

原始XML内容

XML 1

<invoiceCopy>
<document>
    <invoicenumber>1245678</invoicenumber>
    <invoicedate>2024-06-06</invoicedate>
    <pdffilename>12345678.pdf</pdffilename>
    <sendingsystemid>ABC</sendingsystemid>
</document>
</invoicecopy>

XML 2

<invoiceCopy>
<document>
    <invoicenumber>2222222</invoicenumber>
    <invoicedate>2024-06-06</invoicedate>
    <pdffilename>2222222.pdf</pdffilename>
    <sendingsystemid>XYZ</sendingsystemid>
</document>
</invoicecopy>

错误输出

<invoiceCopy>
<document>
    <invoicenumber>1245678</invoicenumber>
    <invoicedate>2024-06-06</invoicedate>
    <pdffilename>12345678.pdf</pdffilename>
    <sendingsystemid>ABC</sendingsystemid>
</document>
<invoiceCopy>
<document>
    <invoicenumber>2222222</invoicenumber>
    <invoicedate>2024-06-06</invoicedate>
    <pdffilename>2222222.pdf</pdffilename>
    <sendingsystemid>XYZ</sendingsystemid>
</document>
</invoicecopy>
</invoicecopy>

运行代码

xml_files = glob.glob(XML_DIR +"/*.xml")
print ("xml files are " , xml_files)
xml_element_tree = None
for filename in xml_files:
    data = ElementTree.parse(filename).getroot()
    print (data)
    for result in data.iter('invoiceCopy'):
     if xml_element_tree is None:
       newfilename = "print" + datetime.today().strftime('%Y%m%d_%H%M%S')+ ".xml"
       print(newfilename)
       xml_element_tree=data
     else:
       for statement in data.iter('invoiceCopy'):
        xml_element_tree.append(statement)

    if xml_element_tree is not None:
    result1 = ''
    m_encoding = "iso-8859-1"
    dom = xml.dom.minidom.parseString(ElementTree.tostring(xml_element_tree))
    xml_string = dom.toprettyxml()
    for result in xml_string.split('\n'):
       if not result.strip() == '':
         result1=result1+result+"\n"
    part1, part2 = result1.split('?&gt;')        
    with open("/data/ebpp/star/fh/invoicecopy/print/" + newfilename, "w") as xfile:
       xfile.write(part1 + 'encoding=\"{}\"?&gt;'.format(m_encoding) + part2)
       xfile.close()

疑问

为何第二个XML会被追加到第一个XML的根节点内部?


原因分析

代码逻辑错误在于:处理第一个文件时,xml_element_tree被赋值为第一个文件的根节点<invoiceCopy>;后续处理第二个文件时,直接调用xml_element_tree.append(statement),将第二个<invoiceCopy>节点作为子节点添加到第一个<invoiceCopy>内部,因此形成嵌套结构。

标准XML要求文档只能有一个顶级根节点,要实现两个<invoiceCopy>节点并列,需创建一个新的容器根节点,将两个<invoiceCopy>都作为子节点添加到该容器下。

修正后的代码

import glob
import datetime
import xml.etree.ElementTree as ElementTree
import xml.dom.minidom

XML_DIR = "你的XML目录路径"  # 替换为实际路径

xml_files = glob.glob(XML_DIR + "/*.xml")
print("xml files are ", xml_files)

# 创建新的容器根节点
root = ElementTree.Element("invoices")
xml_element_tree = ElementTree.ElementTree(root)

for filename in xml_files:
    tree = ElementTree.parse(filename)
    invoice_copy = tree.getroot()
    # 将每个invoiceCopy节点添加到容器根节点下
    root.append(invoice_copy)

# 生成格式化XML并写入文件
if xml_element_tree is not None:
    newfilename = "print" + datetime.today().strftime('%Y%m%d_%H%M%S') + ".xml"
    m_encoding = "iso-8859-1"
    rough_string = ElementTree.tostring(root, 'utf-8')
    dom = xml.dom.minidom.parseString(rough_string)
    xml_string = dom.toprettyxml(encoding=m_encoding)
    
    # 过滤空行
    result1 = ""
    for line in xml_string.split('\n'):
        if line.strip():
            result1 += line + "\n"
    
    # 写入文件(字节流方式避免编码问题)
    with open(f"/data/ebpp/star/fh/invoicecopy/print/{newfilename}", "wb") as xfile:
        xfile.write(result1.encode(m_encoding))

说明

  1. 新增容器根节点:创建<invoices>作为统一顶级节点,符合XML单根节点规范。
  2. 调整节点添加逻辑:将每个文件的<invoiceCopy>节点添加到容器根节点下,而非第一个<invoiceCopy>内部,实现节点并列。
  3. 优化编码与写入:直接使用字节流写入文件,避免字符串拆分拼接导致的编码异常,同时简化空行过滤逻辑。

内容的提问来源于stack exchange,提问作者Pooja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 02:36:04