You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PySpark生成指定结构的List JSON文件?

Solution to Generate Target JSON Structure in PySpark

You're already halfway there with your code to extract the column lists! Now we just need to assemble those lists into the exact JSON structure you want. Here are two straightforward approaches:

Approach 1: Generate JSON String Directly (Python)

If you just need the JSON output as a string (for printing or manual file writing), use Python's built-in json module to construct the structure:

import json

# Your existing code to get column lists
input_data = spark.read.csv("/tmp/test234.csv", header=True, inferSchema=True)
def is_numeric(data_type): return data_type not in ('date', 'string', 'boolean')
def is_nonnumeric(data_type): return data_type in ('string',)  # Fixed tuple syntax
sub = "__"

Loaded_numeric_columns = [name for name, data_type in input_data.dtypes if is_numeric(data_type) and (sub not in name)]
Loaded_category_columns = [name for name, data_type in input_data.dtypes if is_nonnumeric(data_type) and (sub not in name)]
enriched_category_columns = [name for name, data_type in input_data.dtypes if is_nonnumeric(data_type) and (sub in name)]
enriched_index_columns = [name for name, data_type in input_data.dtypes if is_numeric(data_type) and (sub in name)]

# Assemble the target structure
result = [
    {
        "Loaded_data": [
            {
                "Loaded_numeric_columns": Loaded_numeric_columns,
                "Loaded_category_columns": Loaded_category_columns
            }
        ],
        "enriched_data": [
            {
                "enriched_category_columns": enriched_category_columns,
                "enriched_index_columns": enriched_index_columns
            }
        ]
    }
]

# Convert to formatted JSON string
json_output = json.dumps(result, indent=4)
print(json_output)

Approach 2: Write to JSON File Using PySpark

If you need to write the output directly to a JSON file (common in Spark workflows), create a single-row DataFrame matching your structure and write it out:

from pyspark.sql import Row

# Use the same column lists from your existing code

# Create nested Row objects to mirror the JSON structure
loaded_data_row = Row(
    Loaded_numeric_columns=Loaded_numeric_columns,
    Loaded_category_columns=Loaded_category_columns
)

enriched_data_row = Row(
    enriched_category_columns=enriched_category_columns,
    enriched_index_columns=enriched_index_columns
)

final_row = Row(
    Loaded_data=[loaded_data_row],
    enriched_data=[enriched_data_row]
)

# Create a single-row DataFrame
output_df = spark.createDataFrame([final_row])

# Write to JSON (coalesce(1) ensures a single output file, adjust if needed)
output_df.coalesce(1).write.mode("overwrite").json("/path/to/your/output_directory")

Quick Notes:

  • I fixed the is_nonnumeric function to use proper tuple syntax (('string',) instead of ('string') which is just a string).
  • If you're using Python 3, make sure to use parentheses with print() (e.g., print(Loaded_numeric_columns) instead of print Loaded_numeric_columns).
  • When writing with PySpark, the output will be a directory with part files. Using coalesce(1) combines all data into one file, but avoid this for large datasets as it can impact performance.

内容的提问来源于stack exchange,提问作者Shankar Panda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:32:41