You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 3.0中如何将含AM/PM的时间字符串转换为时间戳?

Spark 3.0 时间格式转换报错解决方案

问题场景

原有代码通过unix_timestamp和from_unixtime组合实现时间字符串的解析与格式化,在Spark 3.0中触发以下报错:

org.apache.spark.SparkUpgradeException: You may get a different result due to the upgrading of Spark 3.0: Fail to recognize 'EEE MMM dd HH:mm:ss zzz yyyy' pattern in the DateTimeFormatter.

原代码示例:

from pyspark.sql import Row
df = sc.parallelize([Row(visit_dts='5/1/2018 3:48:14 PM')]).toDF()

import pyspark.sql.functions as f

web = df.withColumn("web_datetime", f.from_unixtime(f.unix_timestamp("visit_dts",'MM/dd/yyyy hh:mm:ss aa'),'MM/dd/yyyy HH:mm:ss'))

预期输出:

+-------------------+-------------------+
|          visit_dts|       web_datetime|
+-------------------+-------------------+
|5/1/2018 3:48:14 PM|05/01/2018 15:48:14|
+-------------------+-------------------+

报错原因

Spark 3.0对时间处理API进行了底层升级:

  1. 原unix_timestamp函数底层实现从SimpleDateFormat切换为Java的DateTimeFormatter,两者格式匹配逻辑存在差异
  2. unix_timestamp在Spark 3.0中已被标记为过时,若使用时隐式依赖默认格式,会触发格式识别失败的异常

解决方案

使用Spark 3.0推荐的to_timestamp+date_format组合替代旧API:

  • to_timestamp:直接将指定格式的字符串解析为Timestamp类型,类型更安全
  • date_format:将Timestamp类型格式化为目标字符串格式,逻辑更直观

修改后代码:

from pyspark.sql import Row
import pyspark.sql.functions as f

df = sc.parallelize([Row(visit_dts='5/1/2018 3:48:14 PM')]).toDF()

web = df.withColumn(
    "web_datetime",
    f.date_format(f.to_timestamp("visit_dts", 'MM/dd/yyyy hh:mm:ss aa'), 'MM/dd/yyyy HH:mm:ss')
)

web.show()

运行后可得到预期输出,且避免了Spark版本升级带来的兼容性问题。

内容的提问来源于stack exchange,提问作者Su1tan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 10:55:21