You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析大型XML仅返回最后一条数据 需导出为pandas DataFrame

问题原因

代码只返回最后一条数据的核心原因是缩进逻辑错误:遍历aixm:DesignatedPointTimeSlice节点的for循环没有包含后续查找坐标、输出结果的逻辑,第一个循环执行完成后,point变量会固定在最后一个匹配到的TimeSlice节点上,后续查找pos、打印的逻辑只会针对这最后一个节点执行,自然拿不到全量数据。
另外代码存在一处不规范写法:查找aixm:type节点时混用了硬编码命名空间和命名空间字典传参的方式,虽然当前没有触发报错,但可维护性很差。

可行修复方案

基础修正版(适配100MB文件,内存占用可控)

直接修正缩进逻辑,统一命名空间写法,遍历每个TimeSlice节点时同步提取对应字段,避免跨循环访问变量:

import xml.etree.ElementTree as ET
import pandas as pd

# 统一命名空间配置
ns = {
    'aixm': 'http://www.aixm.aero/schema/5.1.1',
    'adrext': 'http://www.aixm.aero/schema/5.1.1/extensions/EUR/ADR',
    'gml': 'http://www.opengis.net/gml/3.2'
}

tree = ET.parse('file.xml')
root = tree.getroot()

result = []
# 遍历所有时间片节点,所有字段提取逻辑都放在循环内部
for time_slice in root.findall('.//aixm:DesignatedPointTimeSlice', ns):
    # 提取当前节点下的目标字段
    designator = time_slice.find('.//aixm:designator', ns)
    point_type = time_slice.find('.//aixm:type', ns)
    pos = time_slice.find('.//gml:pos', ns)
    
    # 做空值兼容,避免部分节点缺失字段时触发AttributeError
    data = {
        'designator': designator.text if designator is not None else None,
        'type': point_type.text if point_type is not None else None,
        'pos': pos.text if pos is not None else None
    }
    result.append(data)

# 转为DataFrame,可按需拆分经纬度
df = pd.DataFrame(result)
# 拆分经纬度示例(pos字段为「纬度 经度」空格分隔格式)
df[['lat', 'lon']] = df['pos'].str.split(expand=True).astype(float)
print(df)

大文件优化版(流式解析,内存占用更低)

如果后续XML文件体积继续增大,担心全量加载占用过高内存,可以用iterparse做流式逐节点解析,解析完成就释放无用节点内存,可适配GB级XML文件:

import xml.etree.ElementTree as ET
import pandas as pd

ns = {
    'aixm': 'http://www.aixm.aero/schema/5.1.1',
    'gml': 'http://www.opengis.net/gml/3.2'
}

result = []
# 流式逐事件解析,只监听节点结束事件
for event, elem in ET.iterparse('file.xml', events=('end',)):
    # 匹配到目标TimeSlice节点时提取数据
    if elem.tag == f"{{{ns['aixm']}}}DesignatedPointTimeSlice":
        designator = elem.find('.//aixm:designator', ns)
        point_type = elem.find('.//aixm:type', ns)
        pos = elem.find('.//gml:pos', ns)
        
        result.append({
            'designator': designator.text if designator is not None else None,
            'type': point_type.text if point_type is not None else None,
            'pos': pos.text if pos is not None else None
        })
        # 解析完当前节点后清空子节点,释放内存
        elem.clear()

df = pd.DataFrame(result)
注意事项
  • 所有和单个遍历节点相关的查找、取值逻辑,必须放在对应for循环的缩进块内,不要放到循环外层,否则只会拿到最后一次循环的变量值
  • 提取节点文本前一定要做非空判断,避免部分节点缺失字段时程序中断
  • 100MB的XML用普通parse模式完全可以正常处理,只有文件体积超过500MB时才需要考虑换成iterparse流式方案

内容的提问来源于stack exchange,提问作者setan0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 16:51:19