Python读取XML写入CSV触发IndexError问题求助
问题:Python读取XML转CSV时触发IndexError
尝试用Python脚本读取XML记录并写入CSV,但在获取第一条记录时触发IndexError: list index out of range。
XML文件内容
<records xmlns="http://xmlns.opennms.org/xsd/config/model-import" date-stamp="2025-01-13T11:51:32.229-08:00" foreign-source="Live" last-import="2025-01-13T12:10:16.061-08:00"> <record foreign-id="1734632502410" node-label="A2F0911D_TPMW137A"> <interface ip-addr="11.100.0.59" status="1" snmp-primary="P"> <monitored-service service-name="SNMP"/> <monitored-service service-name="ICMP"/> </interface> <asset name="latitude" value="27.40733333"/> <asset name="longitude" value="-82.53013889"/> <meta-data context="requisition" key="Region" value="SOUTH"/> <meta-data context="requisition" key="Area" value="Florida"/> <meta-data context="requisition" key="Market" value="TP"/> <meta-data context="requisition" key="Cluster" value="A2F0406A"/> <meta-data context="requisition" key="Software version" value="12.1.0.0.0.371"/> <meta-data context="requisition" key="IDU Serial Number" value="E249J10478"/> <meta-data context="requisition" key="Radio slot_sn" value=" Slot 2, Port 2 -RSL=-29.9 E249J10478 /// Slot 2, Port 1 -RSL=-29.9 E249J10478 /// "/> <meta-data context="requisition" key="ComSearch Call Sign" value="WRMD360"/> <meta-data context="requisition" key="Linkids" value="3007262 "/> <meta-data context="requisition" key="Slot_Far_IP" value=" Slot 2, Port 2 11_100_133_187 /// Slot 2, Port 1 11_100_133_187 /// "/> <meta-data context="requisition" key="Crawler First Found Date" value="2024-08-31 00:00:00"/> <meta-data context="requisition" key="Crawler Last Found Date" value="2024-11-29 00:00:00"/> </record> <record foreign-id="1734632502418" node-label="A2F0406A_TPMW136A"> <interface ip-addr="11.100.133.187" status="1" snmp-primary="P"> <monitored-service service-name="SNMP"/> <monitored-service service-name="ICMP"/> </interface> <asset name="latitude" value="27.37944444"/> <asset name="longitude" value="-82.50444444"/> <meta-data context="requisition" key="Region" value="SOUTH"/> <meta-data context="requisition" key="Area" value="Florida"/> <meta-data context="requisition" key="Market" value="TP"/> <meta-data context="requisition" key="Cluster" value="A2F0406A"/> <meta-data context="requisition" key="Software version" value="12.1.0.0.0.371"/> <meta-data context="requisition" key="IDU Serial Number" value="E403L29591"/> <meta-data context="requisition" key="Radio slot_sn" value=" Slot 2, Port 2 -RSL=-29.9 E403L29591 /// Slot 2, Port 1 -RSL=-29.9 E403L29591 /// "/> <meta-data context="requisition" key="ComSearch Call Sign" value="WRMD361"/> <meta-data context="requisition" key="Linkids" value="3007262 "/> <meta-data context="requisition" key="Slot_Far_IP" value=" Slot 2, Port 2 11_100_0_59 /// Slot 2, Port 1 11_100_0_59 /// "/> <meta-data context="requisition" key="Crawler First Found Date" value="2024-08-28 00:00:00"/> <meta-data context="requisition" key="Crawler Last Found Date" value="2024-11-29 00:00:00"/> </record> </records>
Python代码
import xml.etree.ElementTree as ET import os import csv XML_req = 'Live.xml' # path = '/opt/opennms/etc/imports/' # path to current xml requistion... path = '/home/adm_bbyers6/Requisition_import/' full_path = os.path.join(path, XML_req) current_CSV_inventory = 'Microwave_XML_inv.csv' # print('the xml path and file is ', full_path) tree = ET.parse(full_path) root = tree.getroot() # Find all 'record' elements records = root.findall('.//record') fieldnames = [elem.tag for elem in records[0]] # Write to CSV with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() for record in records: row = {child.tag: child.text for child in record} writer.writerow(row)
报错信息
Traceback (most recent call last): File "update_opennms.req.py", line 93, in <module> fieldnames = [elem.tag for elem in records[0]] IndexError: list index out of range
解决方案
问题根源
XML文件带有命名空间xmlns="http://xmlns.opennms.org/xsd/config/model-import",但原代码查找record元素时未指定命名空间,导致records为空列表,访问records[0]触发索引越界。此外,原代码仅提取子元素的标签和文本,未处理XML中存储在属性里的实际数据(如foreign-id、ip-addr等)。
修复后的代码
import xml.etree.ElementTree as ET import os import csv XML_req = 'Live.xml' path = '/home/adm_bbyers6/Requisition_import/' full_path = os.path.join(path, XML_req) # 定义XML命名空间映射 ns = {'nms': 'http://xmlns.opennms.org/xsd/config/model-import'} tree = ET.parse(full_path) root = tree.getroot() # 使用命名空间查找所有record元素 records = root.findall('.//nms:record', ns) if not records: print("未找到任何record元素") exit() # 收集所有可能的字段名 fieldnames = set() for record in records: # 添加record自身的属性字段 fieldnames.update(record.attrib.keys()) # 添加子元素对应的字段 for child in record: if child.tag.endswith('interface'): fieldnames.add('ip-addr') elif child.tag.endswith('asset'): fieldnames.add(child.attrib['name']) elif child.tag.endswith('meta-data'): fieldnames.add(child.attrib['key']) fieldnames = sorted(fieldnames) # 写入CSV文件 with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() for record in records: row = {} # 写入record的属性数据 row.update(record.attrib) # 处理子元素中的数据 for child in record: if child.tag.endswith('interface'): row['ip-addr'] = child.attrib.get('ip-addr') elif child.tag.endswith('asset'): row[child.attrib['name']] = child.attrib.get('value') elif child.tag.endswith('meta-data'): row[child.attrib['key']] = child.attrib.get('value') writer.writerow(row)
代码说明
- 命名空间处理:通过定义命名空间映射
ns,确保能正确定位XML中的record元素。 - 字段收集:遍历所有
record,收集属性、子元素对应的所有字段,避免遗漏数据列。 - 数据提取:分别提取
record的属性、interface的IP地址、asset和meta-data的键值对,构建完整的行数据。 - 空值判断:增加
records为空的判断,提前终止并提示,避免再次触发索引错误。
内容的提问来源于stack exchange,提问作者Brett Byers
相关产品推荐
相关产品推荐

