You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量遍历XML文件解析时出现KeyError: 'occupied'问题求助

遍历文件夹解析XML时触发KeyError: 'occupied',单个文件解析无异常

我是一名编程初学者,现在需要解析一批结构类似下面的XML文件:

<parking id="pucpr">
 <space id="1" occupied="0">
 <rotatedRect>
 <center x="300" y="207" />
 <size w="55" h="32" />
 <angle d="-74" />
 </rotatedRect>
 <contour>
 <point x="278" y="230" />
 <point x="290" y="186" />
 <point x="324" y="185" />
 <point x="308" y="230" />
 </contour>
 </space>
 <space id="2" occupied="0">
 <rotatedRect>
 <center x="332" y="209" />
 <size w="56" h="33" />
 <angle d="-77" />
 </rotatedRect>
 <contour>
 <point x="325" y="185" />
 <point x="355" y="185" />
 <point x="344" y="233" />
 <point x="310" y="233" />
 </contour>
 </space>
...
</parking>

这些文件分布在不同文件夹里,大概有数百个。我写了下面的代码来批量解析:

import xml.etree.ElementTree as ET
import os
import xlsxwriter
data_path = '/Users/jaehyunlee/Desktop/for_test'
# Read full directory and file name in the folder
for path, dirs, files in os.walk(data_path):
    for file in files:
        if os.path.splitext(file)[1].lower() == '.xml': # filtering only for .xml files
            full_path = os.path.join(path, file)
            # Parsing data from .xml file
            tree = ET.parse(full_path)
            root = tree.getroot()
            for space in root.iter('space'):
                car = space.attrib["occupied"]
                car_int = int(car)

现在遇到了一个问题:当我尝试获取<space>标签的occupied属性时,代码返回KeyError: 'occupied'。但解析x、y、w、h这些其他属性都完全正常。而且单独解析一个XML文件的时候不会出现这个错误,只有遍历所有文件批量处理时才会触发这个错误,希望能得到大家的帮助。


问题原因分析

这个错误的核心原因是:你文件夹中的某些XML文件里,存在没有occupied属性的<space>标签。你单独测试的文件里所有<space>都有这个属性,但批量处理时遇到了不符合预期结构的文件/标签,导致直接通过space.attrib["occupied"]取值时触发KeyError。

解决方案

这里有几个可行的处理方式,你可以根据需求选择:

1. 使用字典的get()方法安全取值

dict.get()方法允许你在键不存在时返回一个默认值,避免KeyError:

for space in root.iter('space'):
    # 当没有occupied属性时,默认返回'0'或者你需要的其他值
    car = space.attrib.get("occupied", "0")
    car_int = int(car)

2. 先检查属性是否存在再取值

如果你需要对缺失属性的情况做特殊处理(比如记录日志),可以先判断属性是否存在:

import logging
# 配置日志,方便记录异常文件
logging.basicConfig(filename='xml_parse_errors.log', level=logging.WARNING)

for path, dirs, files in os.walk(data_path):
    for file in files:
        if os.path.splitext(file)[1].lower() == '.xml':
            full_path = os.path.join(path, file)
            tree = ET.parse(full_path)
            root = tree.getroot()
            for space in root.iter('space'):
                if 'occupied' in space.attrib:
                    car = space.attrib["occupied"]
                    car_int = int(car)
                else:
                    # 记录哪个文件的哪个space标签缺失了属性
                    space_id = space.attrib.get('id', 'unknown')
                    logging.warning(f"文件 {full_path} 中的space标签(id: {space_id})缺失occupied属性")
                    # 这里可以给car_int设置默认值,比如0
                    car_int = 0

3. 提前验证所有XML文件的结构

如果你希望确保所有文件都符合预期结构,可以先写一个小脚本扫描所有XML文件,找出缺失occupied属性的文件和标签,手动修复后再批量解析。

额外建议

批量处理文件时,最好加上异常捕获(比如try-except),避免因为单个文件的问题导致整个程序中断:

for path, dirs, files in os.walk(data_path):
    for file in files:
        if os.path.splitext(file)[1].lower() == '.xml':
            full_path = os.path.join(path, file)
            try:
                tree = ET.parse(full_path)
                root = tree.getroot()
                for space in root.iter('space'):
                    car = space.attrib["occupied"]
                    car_int = int(car)
                    # 后续处理逻辑
            except Exception as e:
                logging.error(f"处理文件 {full_path} 时出错: {str(e)}")
                continue

内容的提问来源于stack exchange,提问作者Jaehyun Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:22:40