You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML转CSV及主CSV追加:去除命名空间与流程优化问询

问题解决方案:XML转CSV适配PAD与SQL交互优化

1. 移除表头的XML命名空间(xmlns)

原代码对标签命名空间的处理存在遗漏,导致部分表头仍保留命名空间前缀。通过统一标签清理逻辑解决:

  • 新增clean_tag函数,判断标签是否包含命名空间分隔符},按需截取纯标签名
  • 提前清理所有遍历到的标签,避免循环内重复处理

2. 实现CSV数据追加写入

通过修改to_csv参数实现追加,并根据文件存在状态控制表头写入:

  • 导入os模块检查目标CSV是否存在
  • 写入时设置mode='a'开启追加模式,header=not os.path.exists(路径)避免重复生成表头

3. 适配PAD内置IronPython2.7

IronPython2.7对numpy、pandas支持有限,改用Python标准库替代第三方依赖:

  • 移除numpy、pandas,仅保留xml.etree.ElementTree、csv、os标准库
  • 兼容Python2.7语法(如iteritems()替代items()),手动构建数据并写入CSV

修改后的Python3.10代码(解决前两个问题)

import os
import pandas as pd
from xml.etree import ElementTree

def clean_tag(tag):
    # 清理标签中的命名空间
    if '}' in tag:
        return tag.rsplit('}', 1)[1]
    return tag

maintree = ElementTree.parse('FILE_XML.xml')
parentroot = maintree.getroot()

# 提前清理所有标签的命名空间,去重后作为表头候选
all_tags = list(set([clean_tag(elem.tag) for elem in parentroot.iter()]))

rows = []
for child in parentroot:
    temp_dict = {}
    for inners in child.iter():
        cleaned_tag = clean_tag(inners.tag)
        # 合并属性与标签文本
        temp_dict.update(inners.attrib)
        temp_dict[cleaned_tag] = inners.text
    rows.append(temp_dict)

dataframe = pd.DataFrame.from_dict(rows, orient='columns')
dataframe = dataframe.replace({pd.NA: None})  # 替换空值适配SQL

csv_path = 'FILE_TABLE_CMD.csv'
# 追加写入:文件不存在则创建并写表头,存在则仅追加数据
dataframe.to_csv(csv_path, index=False, mode='a', header=not os.path.exists(csv_path))

IronPython2.7适配代码(兼容PAD内置环境)

import os
import csv
from xml.etree import ElementTree

def clean_tag(tag):
    if '}' in tag:
        return tag.rsplit('}', 1)[1]
    return tag

maintree = ElementTree.parse('FILE_XML.xml')
parentroot = maintree.getroot()

# 收集所有唯一字段(属性键+清理后标签)作为表头
all_fields = set()
for child in parentroot:
    for inners in child.iter():
        all_fields.update(inners.attrib.keys())
        all_fields.add(clean_tag(inners.tag))
headers = list(all_fields)

rows = []
for child in parentroot:
    temp_dict = {}
    for inners in child.iter():
        cleaned_tag = clean_tag(inners.tag)
        # 合并属性
        for k, v in inners.attrib.iteritems():  # Python2.7迭代字典项语法
            temp_dict[k] = v
        # 写入标签文本
        temp_dict[cleaned_tag] = inners.text
    rows.append(temp_dict)

csv_path = 'FILE_TABLE_CMD.csv'
file_exists = os.path.exists(csv_path)

# Python2.7需用二进制模式写入CSV避免换行问题
with open(csv_path, 'ab' if file_exists else 'wb') as csvfile:
    writer = csv.DictWriter(csvfile, fieldnames=headers)
    if not file_exists:
        writer.writeheader()
    # 补全缺失字段,避免写入时KeyError
    for row in rows:
        for header in headers:
            if header not in row:
                row[header] = None
        writer.writerow(row)

内容的提问来源于stack exchange,提问作者Doug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 21:10:36