Spark中无法读取嵌套JSON元素?求解决方法
解决Spark读取嵌套JSON数组时的字段访问报错
报错的核心原因是testPlans.attr3是数组类型,Spark中数组只能通过整数索引(比如attr3[0])访问元素,而你试图用字符串uniqueId当作索引,导致类型不匹配。以下是两种针对性解决方案:
方案一:扁平化数组(适合需要展开数据的场景)
先通过explode函数展开外层的testPlans数组,再处理内层的attr3数组:
方式1:提取attr3数组的第一个元素
import org.apache.spark.sql.functions.explode val df=spark.read.json(Seq(""" [{ "testPlans": [{ "attr1": "abc" }, { "attr2": "bac", "attr3": [{ "uniqueId": "111" }] }] }]""").toDS()) // 展开testPlans数组,将每个数组元素转为单独行 val explodedDf = df.select(explode($"testPlans").alias("testPlan")) // 提取字段,用getItem(0)获取attr3数组的第一个元素,再取uniqueId explodedDf.select( $"testPlan.attr1", $"testPlan.attr2", $"testPlan.attr3.getItem(0).uniqueId".alias("uniqueId") ).show()
方式2:完全展开嵌套数组(如果attr3有多个元素)
如果attr3数组包含多个元素,可以再次使用explode展开:
import org.apache.spark.sql.functions.explode val df=spark.read.json(Seq(""" [{ "testPlans": [{ "attr1": "abc" }, { "attr2": "bac", "attr3": [{ "uniqueId": "111" }, { "uniqueId": "222" }] }] }]""").toDS()) val explodedDf = df.select(explode($"testPlans").alias("testPlan")) // 展开attr3数组,提取每个元素的uniqueId val resultDf = explodedDf.select( $"testPlan.attr1", $"testPlan.attr2", explode($"testPlan.attr3").alias("attr3") ).select( $"attr1", $"attr2", $"attr3.uniqueId" ) resultDf.show()
方案二:保留数组结构(适合需要维持嵌套结构的场景)
使用transform函数遍历attr3数组,提取每个元素的uniqueId,最终得到一个uniqueId的数组:
import org.apache.spark.sql.functions.transform val df=spark.read.json(Seq(""" [{ "testPlans": [{ "attr1": "abc" }, { "attr2": "bac", "attr3": [{ "uniqueId": "111" }, { "uniqueId": "222" }] }] }]""").toDS()) val resultDf = df.select( $"testPlans.attr1", $"testPlans.attr2", transform($"testPlans.attr3", elem => elem.getField("uniqueId")).alias("uniqueIds") ) resultDf.show(truncate = false)
内容的提问来源于stack exchange,提问作者Shankar Panda
相关产品推荐
相关产品推荐

