You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Scala DataFrame中nullable为null的path值存入变量?

提取DataFrame中nullable属性为null的字段对应的path值

需要从解析JSON生成的Spark DataFrame中,筛选出所有**未包含nullable属性或nullable属性值为null**的字段,提取这些字段对应的path值并存入List[String]变量。

完整实现代码

import org.apache.spark.sql.functions._
import spark.implicits._

// 原JSON schema定义
val schema_json = """[{"orders":{"order_id":{"path":"orderid","type":"string"},"customer_id":{"path":"customers.customerId","type":"int","default_value":"null"},"offer_id":{"path":"Offers.Offerid","type":"string"},"eligible":{"path":"eligible.eligiblestatus","type":"string","nullable":true,"default_value":"not eligible"}},"products":{"product_id":{"path":"product_id","type":"string"},"product_name":{"path":"products.productname","type":"string"}}}]"""

// 解析JSON为DataFrame
val schemaRdd = spark.sparkContext.parallelize(schema_json :: Nil)
val schemaRdddf = spark.read.json(schemaRdd)

// 处理orders字段,提取目标path
val targetPaths = schemaRdddf
  // 展开orders下的所有字段为单独行
  .select(explode(map_values($"orders")).alias("field"))
  // 筛选出无nullable属性或nullable值为null的字段
  .filter(!($"field".hasField("nullable")) || $"field.nullable".isNull)
  // 提取path字段值
  .select($"field.path".as[String])
  // 收集为List[String]
  .collect().toList

// 按期望格式输出结果
println(s"abc:List[String]= List(${targetPaths.map(p => s"""\"$p\"""").mkString(",")})")

代码说明

  1. JSON解析:将JSON字符串转为RDD后,通过Spark JSON reader解析成结构化DataFrame。
  2. 字段展开:用map_values提取orders下的所有字段值,再通过explode把每个字段转为独立行,便于逐个处理。
  3. 条件筛选:通过hasField判断字段是否包含nullable属性,或直接判断nullable值是否为null,保留符合要求的字段。
  4. 提取并收集结果:从符合条件的字段中提取path值,最终收集为List[String]类型的变量。

运行结果

执行代码后会输出:

abc:List[String]= List("orderid","customers.customerId","Offers.Offerid")

内容的提问来源于stack exchange,提问作者Shankar Panda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 02:42:43