You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML数据合并:多License合并至单个键的Python代码优化需求

解决XML数据合并:将多个License合并到同一键下

我来帮你搞定这个XML解析的问题!你的需求是把XML中每个区块的Category和对应的多个License合并成单个字典,而不是每个License单独生成一条。下面是修正后的代码,以及对原问题的分析:

修正后的代码

import xml.etree.ElementTree as ET

tree = ET.parse('sample.xml')
root = tree.getroot()

result = []
current_entry = {}

# 遍历根节点下的所有子元素,分组构建条目
for elem in root:
    if elem.tag == 'report_header':
        # 遇到新的服务商头信息,先把之前的条目存入结果(如果存在)
        if current_entry:
            result.append(current_entry)
            current_entry = {}
        # 可选:如果需要保存服务商名称,可以取消下面注释
        # current_entry['provider'] = elem.text.strip()
    elif elem.tag == 'Category':
        # 清理文本中的空格,存入category字段
        current_entry['category'] = elem.text.strip()
    elif elem.tag == 'Licenses':
        # 收集当前区块下的所有License,用逗号连接成字符串
        license_list = [license_elem.text.strip() for license_elem in elem.findall('License')]
        current_entry['License'] = ','.join(license_list)

# 别忘了添加最后一个未处理的条目
if current_entry:
    result.append(current_entry)

# 打印最终结果
for entry in result:
    print(entry)

代码运行结果

{'category': 'Single', 'License': '1234,525'}
{'category': 'Double', 'License': '322,1285,1896'}
{'category': 'Multiple', 'License': '1222'}

原代码的问题分析

  • 标签名称错误:原代码中用findall('./reportheader')查找元素,但XML里的标签是report_header(带有下划线),这会导致找不到任何元素,实际运行会报错。
  • 数据类型误用:way_list被定义为列表,却当作字典来赋值(way_list['category'] = ...),这会触发TypeError。
  • License覆盖问题:原代码每次遍历单个License时都会覆盖字典中的license值,导致每个License单独生成一个字典,而不是合并到同一个条目下。

核心思路说明

  • 遍历XML根节点的所有子元素,每当遇到report_header就开启一个新的条目字典。
  • 遇到Category时,将清理后的文本存入当前条目的category键。
  • 遇到Licenses时,收集所有子节点License的文本,用逗号连接成字符串后存入License键。
  • 每次遇到新的report_header或者遍历结束时,将当前完成的条目存入结果列表。

内容的提问来源于stack exchange,提问作者saikrishna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 08:48:13