如何用Python解析单个文件中的多个XML文档?
解析单文件中多个独立XML片段的方法
你的文件包含多个独立的XML根元素(<infodump>),这不符合XML规范(单个XML文档只能有一个根节点),所以ET.parse()会报错。以下是几种无需拆分文件的解决方法:
方法1:手动分割并逐个解析
读取文件内容后,按<infodump>分割出所有独立片段,再逐个解析:
import xml.etree.ElementTree as ET def parse_multiple_xml(file_path): with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 提取所有完整的<infodump>片段 segments = [f"<infodump>{part}</infodump>" for part in content.split('<infodump>')[1:]] for idx, segment in enumerate(segments): try: root = ET.fromstring(segment) # 这里编写处理单个<infodump>的逻辑,比如提取数据 print(f"处理第{idx+1}个片段,根节点:{root.tag}") for child in root: print(f" {child.tag}: {child.text.strip() if child.text else '无内容'}") except ET.ParseError as e: print(f"解析第{idx+1}个片段出错:{e}") parse_multiple_xml("./id.xml")
方法2:用临时根节点包裹整个内容
给文件内容套一个临时根元素,让它变成合法XML文档后再统一解析:
import xml.etree.ElementTree as ET def parse_with_temp_root(file_path): with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 包裹临时根节点,使内容成为合法XML wrapped_content = f"<temp_root>{content}</temp_root>" root = ET.fromstring(wrapped_content) # 遍历所有<infodump>子节点 for idx, infodump in enumerate(root.findall('infodump')): print(f"处理第{idx+1}个infodump片段") # 处理节点逻辑示例 for child in infodump: print(f" {child.tag}: {child.text.strip() if child.text else '无内容'}") parse_with_temp_root("./id.xml")
方法3:使用SAX解析器(适合大文件)
如果文件体积过大、内存占用敏感,用SAX逐事件解析,捕获每个<infodump>的开始与结束:
from xml.sax import make_parser, handler import xml.etree.ElementTree as ET class InfodumpHandler(handler.ContentHandler): def __init__(self): self.current_segment = "" self.in_infodump = False self.count = 0 def startElement(self, name, attrs): if name == "infodump": self.in_infodump = True self.count += 1 self.current_segment = f"<{name}>" elif self.in_infodump: self.current_segment += f"<{name}>" def characters(self, content): if self.in_infodump: self.current_segment += content def endElement(self, name): if name == "infodump": self.current_segment += f"</{name}>" self.in_infodump = False # 解析并处理当前片段 try: root = ET.fromstring(self.current_segment) print(f"处理第{self.count}个infodump片段") for child in root: print(f" {child.tag}: {child.text.strip() if child.text else '无内容'}") except ET.ParseError as e: print(f"解析第{self.count}个片段出错:{e}") elif self.in_infodump: self.current_segment += f"</{name}>" def parse_large_xml(file_path): parser = make_parser() parser.setContentHandler(InfodumpHandler()) with open(file_path, 'r', encoding='utf-8') as f: parser.parse(f) parse_large_xml("./id.xml")
注意事项
- 若XML片段包含命名空间,需在解析时额外处理(如给标签添加命名空间前缀)
- 方法1、2适合中小文件,方法3适合GB级大文件,避免内存溢出
内容的提问来源于stack exchange,提问作者Don Levey
相关产品推荐
相关产品推荐

