如何在Scala中从Apache Spark Row的Wrapped Array获取List而非Seq
背景
- 从Delta表中以JSON格式获取数据
- 使用Apache Spark和Scala
数据格式
val factories = """ { "cities": { "name": "Sao Paulo", "areas": [ { "code": "41939", "type": "downtown" }, { "code": "48294", "type": "residential" } ] }, "domains": [ { "id": "19sk2nfb", "name" : "defense" } ] } """
代码说明
以下代码从Delta表获取数据并创建样例类对象:fetchedData是按条件筛选出的DataFrame,factoriesSchema是对应的JSON Schema
val structuredData = fetchedData.withColumn( "StructuredFactoryJson", from_json(col("FactoryData"), factoriesSchema) ) val factories = structuredData.collect().map { row => val structJson = row.getAs[Row]("StructuredFactoryJson") val citiesRow = structJson.getAs[Row]("cities") val city = City( citiesRow.getAs[String]("name"), citiesRow .getAs[Seq[Row]]("areas") .map(areaRow => Area( areaRow.getAs[String]("type"), areaRow.getAs[String]("code") ) ) ) val domains = structJson .getAs[Seq[Row]]("domains") .map(domainRow => Domain( domainRow.getAs[String]("id"), domainRow.getAs[String]("name") ) ) Factory(city, domains) }
问题
当前代码能正常运行并得到Seq类型的集合,但希望在保持上层对象构建逻辑不变的前提下,获取List而非Seq,该如何实现?
解决方案
在Scala中直接调用集合的.toList方法,就能把Seq转为List,完全不需要改动上层对象的构建逻辑。
修改后的代码如下:
val factories = structuredData.collect().map { row => val structJson = row.getAs[Row]("StructuredFactoryJson") val citiesRow = structJson.getAs[Row]("cities") val city = City( citiesRow.getAs[String]("name"), // 将areas的Seq转为List citiesRow .getAs[Seq[Row]]("areas") .map(areaRow => Area( areaRow.getAs[String]("type"), areaRow.getAs[String]("code") ) ).toList // 新增转换方法 ) val domains = structJson .getAs[Seq[Row]]("domains") .map(domainRow => Domain( domainRow.getAs[String]("id"), domainRow.getAs[String]("name") ) ).toList // 新增转换方法 Factory(city, domains) }
补充说明
- 如果你的
City、Factory样例类参数本来就定义为List类型,直接加.toList就能完美适配;如果之前是Seq类型,由于List是Seq的子类型,也不需要修改样例类定义。 - 这种方式完全保留了原有的数据转换逻辑,仅在集合转换的最后一步完成类型切换,完全符合需求。
内容的提问来源于stack exchange,提问作者Albatross
相关产品推荐
相关产品推荐

