You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apache Spark 3.x使用from_json解析字符串返回null问题求解

问题原因
  • 类型匹配错误:b列存储的是数组结构的内容,但你传入from_json的schema是单个StructType结构体类型,没有用ArrayType包裹,结构不匹配导致解析失败。
  • 内容格式错误:你构造示例DataFrame时,将原生Python列表对象传入StringType类型的字段,Spark会自动将Python对象转换为其repr表示格式(即{type=abc, unitValue=4.4}这种使用等号、键无引号的格式),而非标准JSON格式,from_json仅支持解析标准JSON字符串,无法识别该格式所以返回全null。
解决方案

场景1:实际业务中b列为标准JSON数组字符串

如果你的生产环境中b列本身就是标准JSON数组格式的字符串,仅需要修改解析用的schema,将结构体用ArrayType包裹即可:

from pyspark.sql.functions import from_json, col
from pyspark.sql.types import StructType, StructField, StringType, ArrayType

# 定义数组类型的schema
json_schema = ArrayType(
    StructType([
        StructField("type", StringType(), nullable=True),
        StructField("unitValue", StringType(), nullable=True)
    ])
)

# 解析得到struct数组类型的新列
df_result = df.withColumn("newvalue", from_json(col("b"), json_schema, {"mode": "PERMISSIVE"}))

场景2:b列为Python对象repr格式的字符串

如果你的数据就是示例中这种非标准的repr格式字符串,需要先做字符串清洗转换为标准JSON,再执行解析:

from pyspark.sql.functions import from_json, col, regexp_replace
from pyspark.sql.types import StructType, StructField, StringType, ArrayType

# 第一步:清洗字符串转成标准JSON格式
df_clean = df.withColumn(
    "b_standard_json",
    # 替换规则:键=值 替换为 "键":值
    regexp_replace(col("b"), r"([\w]+)=", r'"\1":')
)

# 第二步:定义数组类型的schema
json_schema = ArrayType(
    StructType([
        StructField("type", StringType(), nullable=True),
        StructField("unitValue", StringType(), nullable=True)
    ])
)

# 第三步:解析得到struct数组类型的新列
df_result = df_clean.withColumn("newvalue", from_json(col("b_standard_json"), json_schema))

解析后newvalue即为array<struct<type:string, unitValue:string>>类型,可直接通过数组下标+字段名访问内部值,例如col("newvalue")[0].type即可获取第一个元素的type字段值。

内容的提问来源于stack exchange,提问作者Saikat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:30:02