如何使用PySpark生成指定结构的List JSON文件?
Solution to Generate Target JSON Structure in PySpark
You're already halfway there with your code to extract the column lists! Now we just need to assemble those lists into the exact JSON structure you want. Here are two straightforward approaches:
Approach 1: Generate JSON String Directly (Python)
If you just need the JSON output as a string (for printing or manual file writing), use Python's built-in json module to construct the structure:
import json # Your existing code to get column lists input_data = spark.read.csv("/tmp/test234.csv", header=True, inferSchema=True) def is_numeric(data_type): return data_type not in ('date', 'string', 'boolean') def is_nonnumeric(data_type): return data_type in ('string',) # Fixed tuple syntax sub = "__" Loaded_numeric_columns = [name for name, data_type in input_data.dtypes if is_numeric(data_type) and (sub not in name)] Loaded_category_columns = [name for name, data_type in input_data.dtypes if is_nonnumeric(data_type) and (sub not in name)] enriched_category_columns = [name for name, data_type in input_data.dtypes if is_nonnumeric(data_type) and (sub in name)] enriched_index_columns = [name for name, data_type in input_data.dtypes if is_numeric(data_type) and (sub in name)] # Assemble the target structure result = [ { "Loaded_data": [ { "Loaded_numeric_columns": Loaded_numeric_columns, "Loaded_category_columns": Loaded_category_columns } ], "enriched_data": [ { "enriched_category_columns": enriched_category_columns, "enriched_index_columns": enriched_index_columns } ] } ] # Convert to formatted JSON string json_output = json.dumps(result, indent=4) print(json_output)
Approach 2: Write to JSON File Using PySpark
If you need to write the output directly to a JSON file (common in Spark workflows), create a single-row DataFrame matching your structure and write it out:
from pyspark.sql import Row # Use the same column lists from your existing code # Create nested Row objects to mirror the JSON structure loaded_data_row = Row( Loaded_numeric_columns=Loaded_numeric_columns, Loaded_category_columns=Loaded_category_columns ) enriched_data_row = Row( enriched_category_columns=enriched_category_columns, enriched_index_columns=enriched_index_columns ) final_row = Row( Loaded_data=[loaded_data_row], enriched_data=[enriched_data_row] ) # Create a single-row DataFrame output_df = spark.createDataFrame([final_row]) # Write to JSON (coalesce(1) ensures a single output file, adjust if needed) output_df.coalesce(1).write.mode("overwrite").json("/path/to/your/output_directory")
Quick Notes:
- I fixed the
is_nonnumericfunction to use proper tuple syntax (('string',)instead of('string')which is just a string). - If you're using Python 3, make sure to use parentheses with
print()(e.g.,print(Loaded_numeric_columns)instead ofprint Loaded_numeric_columns). - When writing with PySpark, the output will be a directory with part files. Using
coalesce(1)combines all data into one file, but avoid this for large datasets as it can impact performance.
内容的提问来源于stack exchange,提问作者Shankar Panda
相关产品推荐
相关产品推荐

