如何在Python中将指定文本文件转换为嵌套JSON结构?
问题
需要将特定格式的文本文件转换为指定的嵌套JSON结构,便于后续构建表格。
输入文本文件内容
Report_for Reconciliation Execution_of application_1673496470638_0001 Spark_version 2.4.7-amzn-0 Java_version 1.8.0_352 (Amazon.com Inc.) Start_time 2023-01-12 09:45:13.360000 Spark Properties: Job_ID 0 Submission_time 2023-01-12 09:47:20.148000 Run_time 73957ms Result JobSucceeded Number_of_stages 1 Stage_ID 0 Number_of_tasks 16907 Number_of_executed_tasks 16907 Completion_time 73207ms Stage_executed parquet at RawDataPublisher.scala:53 Job_ID 1 Submission_time 2023-01-12 09:48:34.177000 Run_time 11525ms Result JobSucceeded Number_of_stages 2 Stage_ID 1 Number_of_tasks 16907 Number_of_executed_tasks 0 Completion_time 0ms Stage_executed parquet at RawDataPublisher.scala:53 Stage_ID 2 Number_of_tasks 300 Number_of_executed_tasks 300 Completion_time 11520ms Stage_executed parquet at RawDataPublisher.scala:53 Job_ID 2 Submission_time 2023-01-12 09:48:46.908000 Run_time 218358ms Result JobSucceeded Number_of_stages 1 Stage_ID 3 Number_of_tasks 1135 Number_of_executed_tasks 1135 Completion_time 218299ms Stage_executed parquet at RawDataPublisher.scala:53
期望的嵌套JSON结构
{ "Report_for": "Reconciliation", "Execution_of": "application_1673496470638_0001", "Spark_version": "2.4.7-amzn-0", "Java_version": "1.8.0_352 (Amazon.com Inc.)", "Start_time": "2023-01-12 09:45:13.360000", "Job_ID 0": { "Submission_time": "2023-01-12 09:47:20.148000", "Run_time": "73957ms", "Result": "JobSucceeded", "Number_of_stages": "1", "Stage_ID 0": { "Number_of_tasks": "16907", "Number_of_executed_tasks": "16907", "Completion_time": "73207ms", "Stage_executed": "parquet at RawDataPublisher.scala:53", "Stage": "parquet at RawDataPublisher.scala:53" } } }
尝试过的代码及问题
使用defaultdict生成的JSON值为列表形式,无法用于构建表格:
import json from collections import defaultdict INPUT = 'demofile.txt' dict1 = defaultdict(list) def convert(): with open(INPUT) as f: for line in f: command, description = line.strip().split(None, 1) dict1[command].append(description.strip()) OUTPUT = open("demo1file.json", "w") json.dump(dict1, OUTPUT, indent = 4, sort_keys = False)
错误结果片段:
"Report_for": [ "Reconciliation" ], "Execution_of": [ "application_1673496470638_0001" ], "Spark_version": [ "2.4.7-amzn-0" ], "Java_version": [ "1.8.0_352 (Amazon.com Inc.)" ], "Start_time": [ "2023-01-12 09:45:13.360000" ], "Job_ID": [ "0", "1", "2", .... ]]]
解决方案
通过跟踪当前层级(顶级、Job层级、Stage层级),将数据对应到嵌套结构中:
import json INPUT = 'demofile.txt' def convert(): result = {} current_job = None current_stage = None with open(INPUT, 'r') as f: for line in f: line = line.strip() if not line or line == "Spark Properties:": continue parts = line.split(None, 1) if len(parts) < 2: continue key, value = parts[0], parts[1] if key == "Job_ID": job_key = f"Job_ID {value}" result[job_key] = {} current_job = result[job_key] current_stage = None elif key == "Stage_ID": stage_key = f"Stage_ID {value}" current_job[stage_key] = {} current_stage = current_job[stage_key] else: if current_stage is not None: current_stage[key] = value if key == "Stage_executed": current_stage["Stage"] = value elif current_job is not None: current_job[key] = value else: result[key] = value with open("demo1file.json", "w") as output: json.dump(result, output, indent=4, sort_keys=False) if __name__ == "__main__": convert()
逻辑说明
- 用
result存储顶级数据,current_job和current_stage跟踪当前活跃的层级节点。 - 逐行读取文件,跳过空行和无关的
Spark Properties:行。 - 遇到
Job_ID时创建对应Job节点,切换到Job层级;遇到Stage_ID时在当前Job下创建Stage节点,切换到Stage层级。 - 其他字段根据当前层级写入对应节点,同时处理期望中重复的
Stage字段。 - 最后将嵌套结构写入JSON文件。
输出结果示例
生成的完整JSON符合期望格式:
{ "Report_for": "Reconciliation", "Execution_of": "application_1673496470638_0001", "Spark_version": "2.4.7-amzn-0", "Java_version": "1.8.0_352 (Amazon.com Inc.)", "Start_time": "2023-01-12 09:45:13.360000", "Job_ID 0": { "Submission_time": "2023-01-12 09:47:20.148000", "Run_time": "73957ms", "Result": "JobSucceeded", "Number_of_stages": "1", "Stage_ID 0": { "Number_of_tasks": "16907", "Number_of_executed_tasks": "16907", "Completion_time": "73207ms", "Stage_executed": "parquet at RawDataPublisher.scala:53", "Stage": "parquet at RawDataPublisher.scala:53" } }, "Job_ID 1": { "Submission_time": "2023-01-12 09:48:34.177000", "Run_time": "11525ms", "Result": "JobSucceeded", "Number_of_stages": "2", "Stage_ID 1": { "Number_of_tasks": "16907", "Number_of_executed_tasks": "0", "Completion_time": "0ms", "Stage_executed": "parquet at RawDataPublisher.scala:53", "Stage": "parquet at RawDataPublisher.scala:53" }, "Stage_ID 2": { "Number_of_tasks": "300", "Number_of_executed_tasks": "300", "Completion_time": "11520ms", "Stage_executed": "parquet at RawDataPublisher.scala:53", "Stage": "parquet at RawDataPublisher.scala:53" } }, "Job_ID 2": { "Submission_time": "2023-01-12 09:48:46.908000", "Run_time": "218358ms", "Result": "JobSucceeded", "Number_of_stages": "1", "Stage_ID 3": { "Number_of_tasks": "1135", "Number_of_executed_tasks": "1135", "Completion_time": "218299ms", "Stage_executed": "parquet at RawDataPublisher.scala:53", "Stage": "parquet at RawDataPublisher.scala:53" } } }
内容的提问来源于stack exchange,提问作者anonymous
相关产品推荐
相关产品推荐

