Python合并两个XML文件时节点嵌套异常问题求助
合并XML文件时节点嵌套问题排查与解决
问题描述
需合并两个XML文件,使两个<invoiceCopy>节点并列,但运行代码后第二个节点被嵌套在第一个内部。
原始XML内容
XML 1
<invoiceCopy> <document> <invoicenumber>1245678</invoicenumber> <invoicedate>2024-06-06</invoicedate> <pdffilename>12345678.pdf</pdffilename> <sendingsystemid>ABC</sendingsystemid> </document> </invoicecopy>
XML 2
<invoiceCopy> <document> <invoicenumber>2222222</invoicenumber> <invoicedate>2024-06-06</invoicedate> <pdffilename>2222222.pdf</pdffilename> <sendingsystemid>XYZ</sendingsystemid> </document> </invoicecopy>
错误输出
<invoiceCopy> <document> <invoicenumber>1245678</invoicenumber> <invoicedate>2024-06-06</invoicedate> <pdffilename>12345678.pdf</pdffilename> <sendingsystemid>ABC</sendingsystemid> </document> <invoiceCopy> <document> <invoicenumber>2222222</invoicenumber> <invoicedate>2024-06-06</invoicedate> <pdffilename>2222222.pdf</pdffilename> <sendingsystemid>XYZ</sendingsystemid> </document> </invoicecopy> </invoicecopy>
运行代码
xml_files = glob.glob(XML_DIR +"/*.xml") print ("xml files are " , xml_files) xml_element_tree = None for filename in xml_files: data = ElementTree.parse(filename).getroot() print (data) for result in data.iter('invoiceCopy'): if xml_element_tree is None: newfilename = "print" + datetime.today().strftime('%Y%m%d_%H%M%S')+ ".xml" print(newfilename) xml_element_tree=data else: for statement in data.iter('invoiceCopy'): xml_element_tree.append(statement) if xml_element_tree is not None: result1 = '' m_encoding = "iso-8859-1" dom = xml.dom.minidom.parseString(ElementTree.tostring(xml_element_tree)) xml_string = dom.toprettyxml() for result in xml_string.split('\n'): if not result.strip() == '': result1=result1+result+"\n" part1, part2 = result1.split('?>') with open("/data/ebpp/star/fh/invoicecopy/print/" + newfilename, "w") as xfile: xfile.write(part1 + 'encoding=\"{}\"?>'.format(m_encoding) + part2) xfile.close()
疑问
为何第二个XML会被追加到第一个XML的根节点内部?
原因分析
代码逻辑错误在于:处理第一个文件时,xml_element_tree被赋值为第一个文件的根节点<invoiceCopy>;后续处理第二个文件时,直接调用xml_element_tree.append(statement),将第二个<invoiceCopy>节点作为子节点添加到第一个<invoiceCopy>内部,因此形成嵌套结构。
标准XML要求文档只能有一个顶级根节点,要实现两个<invoiceCopy>节点并列,需创建一个新的容器根节点,将两个<invoiceCopy>都作为子节点添加到该容器下。
修正后的代码
import glob import datetime import xml.etree.ElementTree as ElementTree import xml.dom.minidom XML_DIR = "你的XML目录路径" # 替换为实际路径 xml_files = glob.glob(XML_DIR + "/*.xml") print("xml files are ", xml_files) # 创建新的容器根节点 root = ElementTree.Element("invoices") xml_element_tree = ElementTree.ElementTree(root) for filename in xml_files: tree = ElementTree.parse(filename) invoice_copy = tree.getroot() # 将每个invoiceCopy节点添加到容器根节点下 root.append(invoice_copy) # 生成格式化XML并写入文件 if xml_element_tree is not None: newfilename = "print" + datetime.today().strftime('%Y%m%d_%H%M%S') + ".xml" m_encoding = "iso-8859-1" rough_string = ElementTree.tostring(root, 'utf-8') dom = xml.dom.minidom.parseString(rough_string) xml_string = dom.toprettyxml(encoding=m_encoding) # 过滤空行 result1 = "" for line in xml_string.split('\n'): if line.strip(): result1 += line + "\n" # 写入文件(字节流方式避免编码问题) with open(f"/data/ebpp/star/fh/invoicecopy/print/{newfilename}", "wb") as xfile: xfile.write(result1.encode(m_encoding))
说明
- 新增容器根节点:创建
<invoices>作为统一顶级节点,符合XML单根节点规范。 - 调整节点添加逻辑:将每个文件的
<invoiceCopy>节点添加到容器根节点下,而非第一个<invoiceCopy>内部,实现节点并列。 - 优化编码与写入:直接使用字节流写入文件,避免字符串拆分拼接导致的编码异常,同时简化空行过滤逻辑。
内容的提问来源于stack exchange,提问作者Pooja
相关产品推荐
相关产品推荐

