Pyspark无需UDF/RDD:将JSON字符串列表转为字典列表
解决PySpark中JSON字符串数组转结构体列表的问题
问题分析
你的JSON_string列是字符串数组,每个元素是独立的JSON字符串,直接用from_json解析整个数组会失败——因为它不是完整的JSON数组格式,而是由JSON格式的字符串组成的数组。需要逐个解析数组内的每个JSON字符串,再重新组合成结构体列表。
解决方案(无需UDF/RDD)
利用PySpark 3.1+提供的transform函数,对数组中的每个元素单独执行from_json解析,最终返回结构体数组:
定义JSON元素的Schema
由于数组内的JSON结构不完全一致,需包含所有可能的字段并允许为空:from pyspark.sql import functions as F from pyspark.sql.types import StructType, StructField, IntegerType, StringType json_element_schema = StructType([ StructField("Zipcode", IntegerType(), nullable=True), StructField("ZipCodeType", StringType(), nullable=True), StructField("City", StringType(), nullable=True), StructField("State", StringType(), nullable=True) ])转换JSON字符串数组为结构体列表
使用transform遍历数组,对每个字符串执行from_json解析:df_transformed = df.withColumn( "parsed_json_list", F.transform( F.col("JSON_string"), lambda json_str: F.from_json(json_str, json_element_schema) ) )
验证结果
转换后的parsed_json_list列就是你期望的格式:
| Key | parsed_json_list |
|---|---|
| 123456 | [{"Zipcode":704,"ZipCodeType":"STA","City":null,"State":null},{"Zipcode":null,"ZipCodeType":null,"City":"PARC","State":"PR"}] |
| 789123 | [{"Zipcode":7,"ZipCodeType":"AZA","City":null,"State":null},{"Zipcode":null,"ZipCodeType":null,"City":"PRE","State":"XY"}] |
如果不需要null字段,可以后续用dropFields处理,但默认结构已经符合字典列表的要求。
内容的提问来源于stack exchange,提问作者bill
相关产品推荐
相关产品推荐

