Apache Spark 3.x使用from_json解析字符串返回null问题求解
问题原因
- 类型匹配错误:
b列存储的是数组结构的内容,但你传入from_json的schema是单个StructType结构体类型,没有用ArrayType包裹,结构不匹配导致解析失败。 - 内容格式错误:你构造示例DataFrame时,将原生Python列表对象传入
StringType类型的字段,Spark会自动将Python对象转换为其repr表示格式(即{type=abc, unitValue=4.4}这种使用等号、键无引号的格式),而非标准JSON格式,from_json仅支持解析标准JSON字符串,无法识别该格式所以返回全null。
解决方案
场景1:实际业务中b列为标准JSON数组字符串
如果你的生产环境中b列本身就是标准JSON数组格式的字符串,仅需要修改解析用的schema,将结构体用ArrayType包裹即可:
from pyspark.sql.functions import from_json, col from pyspark.sql.types import StructType, StructField, StringType, ArrayType # 定义数组类型的schema json_schema = ArrayType( StructType([ StructField("type", StringType(), nullable=True), StructField("unitValue", StringType(), nullable=True) ]) ) # 解析得到struct数组类型的新列 df_result = df.withColumn("newvalue", from_json(col("b"), json_schema, {"mode": "PERMISSIVE"}))
场景2:b列为Python对象repr格式的字符串
如果你的数据就是示例中这种非标准的repr格式字符串,需要先做字符串清洗转换为标准JSON,再执行解析:
from pyspark.sql.functions import from_json, col, regexp_replace from pyspark.sql.types import StructType, StructField, StringType, ArrayType # 第一步:清洗字符串转成标准JSON格式 df_clean = df.withColumn( "b_standard_json", # 替换规则:键=值 替换为 "键":值 regexp_replace(col("b"), r"([\w]+)=", r'"\1":') ) # 第二步:定义数组类型的schema json_schema = ArrayType( StructType([ StructField("type", StringType(), nullable=True), StructField("unitValue", StringType(), nullable=True) ]) ) # 第三步:解析得到struct数组类型的新列 df_result = df_clean.withColumn("newvalue", from_json(col("b_standard_json"), json_schema))
解析后newvalue即为array<struct<type:string, unitValue:string>>类型,可直接通过数组下标+字段名访问内部值,例如col("newvalue")[0].type即可获取第一个元素的type字段值。
内容的提问来源于stack exchange,提问作者Saikat
相关产品推荐
相关产品推荐

