如何基于Spark DataFrame指定两列的值创建新列?解决结果不符问题
Spark DataFrame 列映射问题解决
原始数据
初始Spark DataFrame数据如下:
client type ------ ---- 89 id 56 id 34 id 13 id 67 phone 68 phone
需求
根据type列的值,把client列的值拆分到新创建的id和phone列:
- 当
type为id时,client值填入id列,phone列留空(null) - 当
type为phone时,client值填入phone列,id列留空(null)
最终要保留所有原始行,预期结果如下:
+--------+----+----+-----+ | client|type| id|phone| +--------+----+----+-----+ | 89| id| 89| null| | 56| id| 56| null| | 34| id| 34| null| | 13| id| 13| null| | 67|phone|null| 67| | 68|phone|null| 68| +--------+----+----+-----+
问题
尝试用这段代码处理:
Df.withColumn("id", when($"type" === "id", $"client")).withColumn("phone", when($"type" === "phone", $"client"))
结果却丢失了type为phone的行,不符合预期。
解决方案
你写的代码逻辑本身没问题,不会导致行丢失,大概率是后续的过滤、显示操作出错了。如果确认是这段代码的问题,可以显式指定otherwise(null)(虽然when默认不满足条件时返回null,但显式声明更清晰),修改后的代码如下:
import org.apache.spark.sql.functions.{when, col} val resultDf = Df .withColumn("id", when(col("type") === "id", col("client")).otherwise(null)) .withColumn("phone", when(col("type") === "phone", col("client")).otherwise(null))
执行这段代码后,就能得到符合预期的结果,所有原始行都会保留,对应列正确填充值或null。
内容的提问来源于stack exchange,提问作者Atum
相关产品推荐
相关产品推荐

