You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark解析含嵌套JSON数组的字符串列并展开遇空值求助

问题原因

你的metrics列不是标准JSON格式:标准JSON要求键必须用双引号包裹,键值对用冒号:分隔,字符串类型的值也要用双引号包裹。但你的数据是{id=1,name=XYZ,value=3}这种格式,导致from_json无法正确解析,最终返回空值。

解决方案

先把非标准字符串转换为标准JSON格式,再进行解析、展开和字段提取:

from pyspark.sql import functions as F
from pyspark.sql.types import ArrayType, StructType, StructField, IntegerType, StringType

# 1. 将非标准字符串转换为标准JSON格式
# 第一步:给键添加双引号,同时把=替换为:
step1_df = df.withColumn(
    "metrics_json",
    F.regexp_replace("metrics", r"(\w+)=", r'"$1":')
)

# 第二步:给字符串类型的值添加双引号(数字值无需处理)
standard_json_df = step1_df.withColumn(
    "metrics_json",
    F.regexp_replace("metrics_json", r":([A-Za-z]+)([,}])", r':"$1"$2')
)

# 2. 手动定义JSON数组的Schema(比自动推断更稳定)
item_schema = StructType([
    StructField("id", IntegerType(), True),
    StructField("name", StringType(), True),
    StructField("value", IntegerType(), True)
])
array_schema = ArrayType(item_schema, True)

# 3. 解析JSON数组并展开为行
parsed_df = standard_json_df.select(
    F.from_json("metrics_json", array_schema).alias("metrics_arr")
)

exploded_df = parsed_df.select(F.explode("metrics_arr").alias("metrics_obj"))

# 4. 提取字段生成最终DataFrame
final_df = exploded_df.select(
    "metrics_obj.id",
    "metrics_obj.name",
    "metrics_obj.value"
)

final_df.show()
预期输出
+---+----+-----+
| id|name|value|
+---+----+-----+
|  1| XYZ|    3|
|  2| KJH|    2|
|  4| ABC|    7|
|  8| HGS|    9|
+---+----+-----+
补充说明

如果name值包含小写字母、下划线等字符,可以把正则表达式调整为:([\w]+)([,}])适配更多场景;若值中存在特殊字符,需根据实际情况修改正则规则。

内容的提问来源于stack exchange,提问作者user10

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 17:10:35