You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析XML输出JSON遇格式及遍历问题求助

解决XML解析转JSON的格式问题与遍历失效问题

问题分析

  1. 遍历失效:parse函数仅处理第一个文件
    你的parse函数在for循环内直接使用return,导致循环第一次迭代就返回结果,后续文件完全没被处理。需要把所有文件的解析结果收集到列表中再统一返回。

  2. JSON格式错误:未用字典构建结构
    当前你把每个字段拼接成带冒号的字符串,转JSON后自然是单独的字符串条目。要生成期望的键值对格式,必须用Python字典来存储每个文件的信息。

修正后的代码

from xml.dom import minidom
import os
import json

# 获取指定目录下的所有文件
def get_files(d):
    return [os.path.join(d, f) for f in os.listdir(d) if os.path.isfile(os.path.join(d,f))]

# 解析XML文件,返回所有文件的结果列表
def parse(files):
    results = []
    for xml_file in files:
        tree = minidom.parse(xml_file)
        
        # 提取字段并构建字典
        trial_info = {
            "NCT ID": tree.getElementsByTagName("nct_id")[0].firstChild.data,
            "brief title": tree.getElementsByTagName("brief_title")[0].firstChild.data,
            "official title": tree.getElementsByTagName("official_title")[0].firstChild.data
        }
        results.append(trial_info)
    return results

# 输出JSON格式结果
def printjson(results):
    for result in results:
        # 增加indent参数让JSON结构更易读
        output_json = json.dumps(result, indent=4)
        print(output_json)

# 执行流程
printjson(parse(get_files('my files path')))

修正后的输出

{
    "NCT ID": "NCT00571389",
    "brief title": "Isolation and Culture of Immune Cells and Circulating Tumor Cells From Peripheral Blood and Leukapheresis Products",
    "official title": "A Study to Facilitate Development of an Ex-Vivo Device Platform for Circulating Tumor Cell and Immune Cell Harvesting, Banking, and Apoptosis-Viability Assay"
}

关键修改点

  • 在parse函数中初始化空列表results,循环内将每个文件的字典信息添加进去,最后返回整个列表,确保所有文件都被处理。
  • 用字典trial_info存储字段的键值对,替代原有的字符串拼接,让json.dumps能生成标准的JSON对象格式。
  • 给json.dumps添加indent=4参数,让输出的JSON结构和你的期望格式一致,更易阅读。

内容的提问来源于stack exchange,提问作者Quan Hoang Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:40:22