You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将指定文本文件转换为嵌套JSON结构?

问题

需要将特定格式的文本文件转换为指定的嵌套JSON结构,便于后续构建表格。

输入文本文件内容

Report_for Reconciliation
Execution_of application_1673496470638_0001
Spark_version 2.4.7-amzn-0
Java_version 1.8.0_352 (Amazon.com Inc.)
Start_time 2023-01-12 09:45:13.360000
Spark Properties: 
Job_ID 0
Submission_time 2023-01-12 09:47:20.148000
Run_time 73957ms
Result JobSucceeded
Number_of_stages 1
Stage_ID 0
Number_of_tasks 16907
Number_of_executed_tasks 16907
Completion_time 73207ms
Stage_executed parquet at RawDataPublisher.scala:53
Job_ID 1
Submission_time 2023-01-12 09:48:34.177000
Run_time 11525ms
Result JobSucceeded
Number_of_stages 2
Stage_ID 1
Number_of_tasks 16907
Number_of_executed_tasks 0
Completion_time 0ms
Stage_executed parquet at RawDataPublisher.scala:53
Stage_ID 2
Number_of_tasks 300
Number_of_executed_tasks 300
Completion_time 11520ms
Stage_executed parquet at RawDataPublisher.scala:53
Job_ID 2
Submission_time 2023-01-12 09:48:46.908000
Run_time 218358ms
Result JobSucceeded
Number_of_stages 1
Stage_ID 3
Number_of_tasks 1135
Number_of_executed_tasks 1135
Completion_time 218299ms
Stage_executed parquet at RawDataPublisher.scala:53

期望的嵌套JSON结构

{
    "Report_for": "Reconciliation",
    "Execution_of": "application_1673496470638_0001",
    "Spark_version": "2.4.7-amzn-0",
    "Java_version": "1.8.0_352 (Amazon.com Inc.)",
    "Start_time": "2023-01-12 09:45:13.360000",
    "Job_ID 0": {
        "Submission_time": "2023-01-12 09:47:20.148000",
        "Run_time": "73957ms",
        "Result": "JobSucceeded",
        "Number_of_stages": "1",
        "Stage_ID 0": {
            "Number_of_tasks": "16907",
            "Number_of_executed_tasks": "16907",
            "Completion_time": "73207ms",
            "Stage_executed": "parquet at RawDataPublisher.scala:53",
            "Stage": "parquet at RawDataPublisher.scala:53"
         }
     }
}

尝试过的代码及问题

使用defaultdict生成的JSON值为列表形式,无法用于构建表格:

import json
from collections import defaultdict

INPUT = 'demofile.txt'
dict1 = defaultdict(list)

def convert():
    with open(INPUT) as f:
        for line in f:
            command, description = line.strip().split(None, 1)
            dict1[command].append(description.strip())
    OUTPUT = open("demo1file.json", "w")
    json.dump(dict1, OUTPUT, indent = 4, sort_keys = False)

错误结果片段:

"Report_for": [ "Reconciliation" ], 
     "Execution_of": [ "application_1673496470638_0001" ], 
     "Spark_version": [ "2.4.7-amzn-0" ], 
     "Java_version": [ "1.8.0_352 (Amazon.com Inc.)" ], 
     "Start_time": [ "2023-01-12 09:45:13.360000" ], 
      "Job_ID": [ 
           "0", 
           "1", 
           "2", ....
]]]

解决方案

通过跟踪当前层级(顶级、Job层级、Stage层级),将数据对应到嵌套结构中:

import json

INPUT = 'demofile.txt'

def convert():
    result = {}
    current_job = None
    current_stage = None
    
    with open(INPUT, 'r') as f:
        for line in f:
            line = line.strip()
            if not line or line == "Spark Properties:":
                continue
            
            parts = line.split(None, 1)
            if len(parts) < 2:
                continue
            key, value = parts[0], parts[1]
            
            if key == "Job_ID":
                job_key = f"Job_ID {value}"
                result[job_key] = {}
                current_job = result[job_key]
                current_stage = None
            elif key == "Stage_ID":
                stage_key = f"Stage_ID {value}"
                current_job[stage_key] = {}
                current_stage = current_job[stage_key]
            else:
                if current_stage is not None:
                    current_stage[key] = value
                    if key == "Stage_executed":
                        current_stage["Stage"] = value
                elif current_job is not None:
                    current_job[key] = value
                else:
                    result[key] = value
    
    with open("demo1file.json", "w") as output:
        json.dump(result, output, indent=4, sort_keys=False)

if __name__ == "__main__":
    convert()

逻辑说明

  1. 用result存储顶级数据,current_job和current_stage跟踪当前活跃的层级节点。
  2. 逐行读取文件,跳过空行和无关的Spark Properties:行。
  3. 遇到Job_ID时创建对应Job节点,切换到Job层级;遇到Stage_ID时在当前Job下创建Stage节点,切换到Stage层级。
  4. 其他字段根据当前层级写入对应节点,同时处理期望中重复的Stage字段。
  5. 最后将嵌套结构写入JSON文件。

输出结果示例

生成的完整JSON符合期望格式:

{
    "Report_for": "Reconciliation",
    "Execution_of": "application_1673496470638_0001",
    "Spark_version": "2.4.7-amzn-0",
    "Java_version": "1.8.0_352 (Amazon.com Inc.)",
    "Start_time": "2023-01-12 09:45:13.360000",
    "Job_ID 0": {
        "Submission_time": "2023-01-12 09:47:20.148000",
        "Run_time": "73957ms",
        "Result": "JobSucceeded",
        "Number_of_stages": "1",
        "Stage_ID 0": {
            "Number_of_tasks": "16907",
            "Number_of_executed_tasks": "16907",
            "Completion_time": "73207ms",
            "Stage_executed": "parquet at RawDataPublisher.scala:53",
            "Stage": "parquet at RawDataPublisher.scala:53"
        }
    },
    "Job_ID 1": {
        "Submission_time": "2023-01-12 09:48:34.177000",
        "Run_time": "11525ms",
        "Result": "JobSucceeded",
        "Number_of_stages": "2",
        "Stage_ID 1": {
            "Number_of_tasks": "16907",
            "Number_of_executed_tasks": "0",
            "Completion_time": "0ms",
            "Stage_executed": "parquet at RawDataPublisher.scala:53",
            "Stage": "parquet at RawDataPublisher.scala:53"
        },
        "Stage_ID 2": {
            "Number_of_tasks": "300",
            "Number_of_executed_tasks": "300",
            "Completion_time": "11520ms",
            "Stage_executed": "parquet at RawDataPublisher.scala:53",
            "Stage": "parquet at RawDataPublisher.scala:53"
        }
    },
    "Job_ID 2": {
        "Submission_time": "2023-01-12 09:48:46.908000",
        "Run_time": "218358ms",
        "Result": "JobSucceeded",
        "Number_of_stages": "1",
        "Stage_ID 3": {
            "Number_of_tasks": "1135",
            "Number_of_executed_tasks": "1135",
            "Completion_time": "218299ms",
            "Stage_executed": "parquet at RawDataPublisher.scala:53",
            "Stage": "parquet at RawDataPublisher.scala:53"
        }
    }
}

内容的提问来源于stack exchange,提问作者anonymous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 05:41:18