You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pyspark无需UDF/RDD:将JSON字符串列表转为字典列表

解决PySpark中JSON字符串数组转结构体列表的问题

问题分析

你的JSON_string列是字符串数组,每个元素是独立的JSON字符串,直接用from_json解析整个数组会失败——因为它不是完整的JSON数组格式,而是由JSON格式的字符串组成的数组。需要逐个解析数组内的每个JSON字符串,再重新组合成结构体列表。

解决方案(无需UDF/RDD)

利用PySpark 3.1+提供的transform函数,对数组中的每个元素单独执行from_json解析,最终返回结构体数组:

  1. 定义JSON元素的Schema
    由于数组内的JSON结构不完全一致,需包含所有可能的字段并允许为空:

    from pyspark.sql import functions as F
    from pyspark.sql.types import StructType, StructField, IntegerType, StringType
    
    json_element_schema = StructType([
        StructField("Zipcode", IntegerType(), nullable=True),
        StructField("ZipCodeType", StringType(), nullable=True),
        StructField("City", StringType(), nullable=True),
        StructField("State", StringType(), nullable=True)
    ])
    
  2. 转换JSON字符串数组为结构体列表
    使用transform遍历数组,对每个字符串执行from_json解析:

    df_transformed = df.withColumn(
        "parsed_json_list",
        F.transform(
            F.col("JSON_string"),
            lambda json_str: F.from_json(json_str, json_element_schema)
        )
    )
    

验证结果

转换后的parsed_json_list列就是你期望的格式:

Keyparsed_json_list
123456[{"Zipcode":704,"ZipCodeType":"STA","City":null,"State":null},{"Zipcode":null,"ZipCodeType":null,"City":"PARC","State":"PR"}]
789123[{"Zipcode":7,"ZipCodeType":"AZA","City":null,"State":null},{"Zipcode":null,"ZipCodeType":null,"City":"PRE","State":"XY"}]

如果不需要null字段,可以后续用dropFields处理,但默认结构已经符合字典列表的要求。

内容的提问来源于stack exchange,提问作者bill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 13:11:13