You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取嵌套XML标签导出CSV遇阻:无法获取地址嵌套字段

解决Python读取XML嵌套标签并导出CSV的问题

问题说明

使用Python读取嵌套XML标签并保存为CSV时,无法读取ADDRESS标签下嵌套的STREET和CITY数据,仅能提取PLANT标签下的外层字段(COMMON、BOTANICAL、ZONE等),需要导出包含Street、City列的完整CSV表格。

测试XML文件

<?xml version="1.0" encoding="UTF-8"?>

<CATALOG>

  <PLANT>

    <COMMON>Bloodroot</COMMON>

    <BOTANICAL>Sanguinaria canadensis</BOTANICAL>

    <ZONE>4</ZONE>

    <LIGHT>Mostly Shady</LIGHT>

    <PRICE>$2.44</PRICE>

    <ADDRESS>
         <STREET>1</STREET>
         <CITY>toronto</CITY>
    </ADDRESS>

    <AVAILABILITY>031599</AVAILABILITY>

  </PLANT>

  <PLANT>

    <COMMON>Columbine</COMMON>

    <BOTANICAL>Aquilegia canadensis</BOTANICAL>

    <ZONE>3</ZONE>

    <LIGHT>Mostly Shady</LIGHT>

    <PRICE>$9.37</PRICE>

    <ADDRESS>
         <STREET>2</STREET>
         <CITY>montreal</CITY>
    </ADDRESS>

    <AVAILABILITY>030699</AVAILABILITY>

  </PLANT>

</CATALOG>

现有代码问题

现有代码直接通过plant.find("STREET")和plant.find("CITY")查找嵌套标签,但STREET和CITY是ADDRESS的子节点,并非PLANT的直接子节点,因此会返回None,访问.text时会抛出AttributeError。此外代码中if(plant)后缺少冒号,存在语法错误。

修复后的代码

from xml.etree import ElementTree
import csv

# 解析XML文件
xml = ElementTree.parse("/dbfs/mnt/ods-outbound/xml_test/plant_catalog.xml")

# 打开CSV文件准备写入(使用with语句自动管理文件关闭)
with open("/dbfs/mnt/ods-outbound/xml_test/plant_catalog.csv", 'w', encoding='utf-8', newline='') as csvfile:
    csvfile_writer = csv.writer(csvfile)
    # 写入表头
    csvfile_writer.writerow(["common", "botanical", "zone", "light", "price", "availability", "street", "city"])
    
    # 遍历每个PLANT节点
    for plant in xml.findall("PLANT"):
        # 提取外层字段
        common = plant.find("COMMON")
        botanical = plant.find("BOTANICAL")
        zone = plant.find("ZONE")
        light = plant.find("LIGHT")
        price = plant.find("PRICE")
        availability = plant.find("AVAILABILITY")
        
        # 先获取ADDRESS节点,再提取嵌套的STREET和CITY
        address = plant.find("ADDRESS")
        street = address.find("STREET") if address is not None else None
        city = address.find("CITY") if address is not None else None
        
        # 处理字段值,避免None导致的AttributeError
        csv_line = [
            common.text if common is not None else "",
            botanical.text if botanical is not None else "",
            zone.text if zone is not None else "",
            light.text if light is not None else "",
            price.text if price is not None else "",
            availability.text if availability is not None else "",
            street.text if street is not None else "",
            city.text if city is not None else ""
        ]
        
        # 写入CSV行
        csvfile_writer.writerow(csv_line)

关键修改点

  1. 嵌套节点查找:先通过plant.find("ADDRESS")获取ADDRESS节点,再从该节点下查找STREET和CITY,确保路径正确。
  2. 空值处理:对所有字段添加判空逻辑,避免因节点不存在导致的AttributeError。
  3. 语法与规范修正:修复if(plant)缺少冒号的问题,同时使用with语句管理文件,自动处理文件关闭,更安全规范。

内容的提问来源于stack exchange,提问作者sajid156

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:55:17