XML转CSV及主CSV追加:去除命名空间与流程优化问询
问题解决方案:XML转CSV适配PAD与SQL交互优化
1. 移除表头的XML命名空间(xmlns)
原代码对标签命名空间的处理存在遗漏,导致部分表头仍保留命名空间前缀。通过统一标签清理逻辑解决:
- 新增
clean_tag函数,判断标签是否包含命名空间分隔符},按需截取纯标签名 - 提前清理所有遍历到的标签,避免循环内重复处理
2. 实现CSV数据追加写入
通过修改to_csv参数实现追加,并根据文件存在状态控制表头写入:
- 导入
os模块检查目标CSV是否存在 - 写入时设置
mode='a'开启追加模式,header=not os.path.exists(路径)避免重复生成表头
3. 适配PAD内置IronPython2.7
IronPython2.7对numpy、pandas支持有限,改用Python标准库替代第三方依赖:
- 移除numpy、pandas,仅保留
xml.etree.ElementTree、csv、os标准库 - 兼容Python2.7语法(如
iteritems()替代items()),手动构建数据并写入CSV
修改后的Python3.10代码(解决前两个问题)
import os import pandas as pd from xml.etree import ElementTree def clean_tag(tag): # 清理标签中的命名空间 if '}' in tag: return tag.rsplit('}', 1)[1] return tag maintree = ElementTree.parse('FILE_XML.xml') parentroot = maintree.getroot() # 提前清理所有标签的命名空间,去重后作为表头候选 all_tags = list(set([clean_tag(elem.tag) for elem in parentroot.iter()])) rows = [] for child in parentroot: temp_dict = {} for inners in child.iter(): cleaned_tag = clean_tag(inners.tag) # 合并属性与标签文本 temp_dict.update(inners.attrib) temp_dict[cleaned_tag] = inners.text rows.append(temp_dict) dataframe = pd.DataFrame.from_dict(rows, orient='columns') dataframe = dataframe.replace({pd.NA: None}) # 替换空值适配SQL csv_path = 'FILE_TABLE_CMD.csv' # 追加写入:文件不存在则创建并写表头,存在则仅追加数据 dataframe.to_csv(csv_path, index=False, mode='a', header=not os.path.exists(csv_path))
IronPython2.7适配代码(兼容PAD内置环境)
import os import csv from xml.etree import ElementTree def clean_tag(tag): if '}' in tag: return tag.rsplit('}', 1)[1] return tag maintree = ElementTree.parse('FILE_XML.xml') parentroot = maintree.getroot() # 收集所有唯一字段(属性键+清理后标签)作为表头 all_fields = set() for child in parentroot: for inners in child.iter(): all_fields.update(inners.attrib.keys()) all_fields.add(clean_tag(inners.tag)) headers = list(all_fields) rows = [] for child in parentroot: temp_dict = {} for inners in child.iter(): cleaned_tag = clean_tag(inners.tag) # 合并属性 for k, v in inners.attrib.iteritems(): # Python2.7迭代字典项语法 temp_dict[k] = v # 写入标签文本 temp_dict[cleaned_tag] = inners.text rows.append(temp_dict) csv_path = 'FILE_TABLE_CMD.csv' file_exists = os.path.exists(csv_path) # Python2.7需用二进制模式写入CSV避免换行问题 with open(csv_path, 'ab' if file_exists else 'wb') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=headers) if not file_exists: writer.writeheader() # 补全缺失字段,避免写入时KeyError for row in rows: for header in headers: if header not in row: row[header] = None writer.writerow(row)
内容的提问来源于stack exchange,提问作者Doug
相关产品推荐
相关产品推荐

