Python解析XML输出JSON遇格式及遍历问题求助
解决XML解析转JSON的格式问题与遍历失效问题
问题分析
遍历失效:parse函数仅处理第一个文件
你的parse函数在for循环内直接使用return,导致循环第一次迭代就返回结果,后续文件完全没被处理。需要把所有文件的解析结果收集到列表中再统一返回。JSON格式错误:未用字典构建结构
当前你把每个字段拼接成带冒号的字符串,转JSON后自然是单独的字符串条目。要生成期望的键值对格式,必须用Python字典来存储每个文件的信息。
修正后的代码
from xml.dom import minidom import os import json # 获取指定目录下的所有文件 def get_files(d): return [os.path.join(d, f) for f in os.listdir(d) if os.path.isfile(os.path.join(d,f))] # 解析XML文件,返回所有文件的结果列表 def parse(files): results = [] for xml_file in files: tree = minidom.parse(xml_file) # 提取字段并构建字典 trial_info = { "NCT ID": tree.getElementsByTagName("nct_id")[0].firstChild.data, "brief title": tree.getElementsByTagName("brief_title")[0].firstChild.data, "official title": tree.getElementsByTagName("official_title")[0].firstChild.data } results.append(trial_info) return results # 输出JSON格式结果 def printjson(results): for result in results: # 增加indent参数让JSON结构更易读 output_json = json.dumps(result, indent=4) print(output_json) # 执行流程 printjson(parse(get_files('my files path')))
修正后的输出
{ "NCT ID": "NCT00571389", "brief title": "Isolation and Culture of Immune Cells and Circulating Tumor Cells From Peripheral Blood and Leukapheresis Products", "official title": "A Study to Facilitate Development of an Ex-Vivo Device Platform for Circulating Tumor Cell and Immune Cell Harvesting, Banking, and Apoptosis-Viability Assay" }
关键修改点
- 在
parse函数中初始化空列表results,循环内将每个文件的字典信息添加进去,最后返回整个列表,确保所有文件都被处理。 - 用字典
trial_info存储字段的键值对,替代原有的字符串拼接,让json.dumps能生成标准的JSON对象格式。 - 给
json.dumps添加indent=4参数,让输出的JSON结构和你的期望格式一致,更易阅读。
内容的提问来源于stack exchange,提问作者Quan Hoang Nguyen
相关产品推荐
相关产品推荐

