You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark DataFrame(Scala)中如何转换为TimestampType避免空值

解决Spark中自定义时间格式转TimestampType为空的问题

直接用cast(TimestampType)转换"11/14/2022 4:48:24 PM"这类格式会失败,因为Spark默认只识别ISO标准的时间格式(比如yyyy-MM-dd HH:mm:ss),无法解析美式日期+12小时制的时间字符串。你需要用to_timestamp函数显式指定匹配的格式:

修改后的代码示例

import org.apache.spark.sql.functions.to_timestamp

val messages = df.withColumn("Offset", $"Offset".cast(LongType))
  .withColumn("Time(readable)", to_timestamp($"EnqueuedTimeUtc", "MM/dd/yyyy h:mm:ss a"))
  .withColumn("Body", $"Body".cast(StringType))
  .select("Offset", "Time(readable)", "Body")

额外处理转换失败的场景

如果存在格式不匹配的异常数据,不想直接得到null,可以用以下两种方式:

  • 用coalesce保留原始字符串或设置默认值:
import org.apache.spark.sql.functions.{to_timestamp, coalesce, lit}

val messages = df.withColumn("Offset", $"Offset".cast(LongType))
  .withColumn("Time(readable)", coalesce(to_timestamp($"EnqueuedTimeUtc", "MM/dd/yyyy h:mm:ss a"), lit("1970-01-01 00:00:00").cast(TimestampType)))
  .withColumn("Body", $"Body".cast(StringType))
  .select("Offset", "Time(readable)", "Body")
  • 用Spark 3.0+支持的try_cast函数,转换失败时返回null但不会抛出错误:
import org.apache.spark.sql.functions.try_cast

val messages = df.withColumn("Offset", $"Offset".cast(LongType))
  .withColumn("Time(readable)", try_cast($"EnqueuedTimeUtc", "timestamp", "MM/dd/yyyy h:mm:ss a"))
  .withColumn("Body", $"Body".cast(StringType))
  .select("Offset", "Time(readable)", "Body")

内容的提问来源于stack exchange,提问作者SONIA_29

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 18:05:43