You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析单个文件中的多个XML文档?

解析单文件中多个独立XML片段的方法

你的文件包含多个独立的XML根元素(<infodump>),这不符合XML规范(单个XML文档只能有一个根节点),所以ET.parse()会报错。以下是几种无需拆分文件的解决方法:

方法1:手动分割并逐个解析

读取文件内容后,按<infodump>分割出所有独立片段,再逐个解析:

import xml.etree.ElementTree as ET

def parse_multiple_xml(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read()
    
    # 提取所有完整的<infodump>片段
    segments = [f"<infodump>{part}</infodump>" for part in content.split('<infodump>')[1:]]
    
    for idx, segment in enumerate(segments):
        try:
            root = ET.fromstring(segment)
            # 这里编写处理单个<infodump>的逻辑,比如提取数据
            print(f"处理第{idx+1}个片段,根节点:{root.tag}")
            for child in root:
                print(f"  {child.tag}: {child.text.strip() if child.text else '无内容'}")
        except ET.ParseError as e:
            print(f"解析第{idx+1}个片段出错:{e}")

parse_multiple_xml("./id.xml")

方法2:用临时根节点包裹整个内容

给文件内容套一个临时根元素,让它变成合法XML文档后再统一解析:

import xml.etree.ElementTree as ET

def parse_with_temp_root(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read()
    
    # 包裹临时根节点,使内容成为合法XML
    wrapped_content = f"<temp_root>{content}</temp_root>"
    root = ET.fromstring(wrapped_content)
    
    # 遍历所有<infodump>子节点
    for idx, infodump in enumerate(root.findall('infodump')):
        print(f"处理第{idx+1}个infodump片段")
        # 处理节点逻辑示例
        for child in infodump:
            print(f"  {child.tag}: {child.text.strip() if child.text else '无内容'}")

parse_with_temp_root("./id.xml")

方法3:使用SAX解析器(适合大文件)

如果文件体积过大、内存占用敏感,用SAX逐事件解析,捕获每个<infodump>的开始与结束:

from xml.sax import make_parser, handler
import xml.etree.ElementTree as ET

class InfodumpHandler(handler.ContentHandler):
    def __init__(self):
        self.current_segment = ""
        self.in_infodump = False
        self.count = 0
    
    def startElement(self, name, attrs):
        if name == "infodump":
            self.in_infodump = True
            self.count += 1
            self.current_segment = f"<{name}>"
        elif self.in_infodump:
            self.current_segment += f"<{name}>"
    
    def characters(self, content):
        if self.in_infodump:
            self.current_segment += content
    
    def endElement(self, name):
        if name == "infodump":
            self.current_segment += f"</{name}>"
            self.in_infodump = False
            # 解析并处理当前片段
            try:
                root = ET.fromstring(self.current_segment)
                print(f"处理第{self.count}个infodump片段")
                for child in root:
                    print(f"  {child.tag}: {child.text.strip() if child.text else '无内容'}")
            except ET.ParseError as e:
                print(f"解析第{self.count}个片段出错:{e}")
        elif self.in_infodump:
            self.current_segment += f"</{name}>"

def parse_large_xml(file_path):
    parser = make_parser()
    parser.setContentHandler(InfodumpHandler())
    with open(file_path, 'r', encoding='utf-8') as f:
        parser.parse(f)

parse_large_xml("./id.xml")

注意事项

  • 若XML片段包含命名空间,需在解析时额外处理(如给标签添加命名空间前缀)
  • 方法1、2适合中小文件,方法3适合GB级大文件,避免内存溢出

内容的提问来源于stack exchange,提问作者Don Levey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:15:37