如何在Apache Spark中转换数组内Struct字段y.c为Timestamp类型?
将Spark DataFrame数组中Struct的指定字段转换为Timestamp类型
问题场景
给定以下Spark DataFrame,其中y字段是包含Struct元素的可选数组,Struct内的c字段为字符串格式的时间数据:
import org.apache.spark.sql.functions._ import org.apache.spark.sql.types._ case class Rec3(i: Long, j: Boolean) case class Rec2(a: Int, b:Rec3, c: String, d: Int) case class Rec1(x:Int, y:Option[Seq[Rec2]], z:Boolean, zz: String) val df = Seq( Rec1(5, Some(Seq(Rec2(4, Rec3(3L, true), "2022-09-22 13:00:00", 3), Rec2(44, Rec3(33L, true), "2022-11-11 22:11:00", 3))), false, "2022-09-23 14:30:00"), Rec1(55, Some(Seq(Rec2(44, Rec3(33L, false), "2023-01-11 21:00:00", 33))), true, "2023-01-22 11:33:00") ).toDF
需求是仅将y数组内每个Struct的c字段转换为Timestamp类型,原尝试代码因类型不匹配报错:
df.withColumn( "y", transform( col("y"), elem => elem.withField( "c", unix_timestamp( col("y.c"), "yyyy-MM-dd' 'HH:mm:ss" ) ) ) )
错误原因
col("y.c")会返回整个数组中所有元素的c字段组成的数组类型,而unix_timestamp函数要求输入单个字符串/日期类型参数,因此触发类型不匹配错误。
解决方案
在transform的迭代逻辑中,通过当前元素elem直接访问Struct内的c字段,使用to_timestamp函数完成类型转换(该函数直接返回Timestamp类型,比unix_timestamp更贴合需求):
val resultDf = df.withColumn( "y", transform( col("y"), elem => elem.withField( "c", to_timestamp(elem.c, "yyyy-MM-dd HH:mm:ss") ) ) )
验证结果
执行resultDf.printSchema()可以看到y数组内的c字段已变为Timestamp类型:
root |-- x: integer (nullable = false) |-- y: array (nullable = true) | |-- element: struct (containsNull = true) | | |-- a: integer (nullable = false) | | |-- b: struct (nullable = false) | | | |-- i: long (nullable = false) | | | |-- j: boolean (nullable = false) | | |-- c: timestamp (nullable = true) | | |-- d: integer (nullable = false) |-- z: boolean (nullable = false) |-- zz: string (nullable = true)
说明
transform函数遍历数组中的每个Struct元素elem,elem.c直接指向当前元素的c字段(单个字符串值),符合to_timestamp的参数要求。- 原字段
y是Option[Seq[Rec2]],transform会自动处理None的情况,转换后仍保持Option类型。
内容的提问来源于stack exchange,提问作者54l3d
相关产品推荐
相关产品推荐

