Apache Spark中如何将澳大利亚/墨尔本时区的日期时间转换为UTC?
问题解答:Spark Scala下墨尔本时区日期字符串转UTC时间
以下为Spark开发场景下两种最常用的实现方案,原始输入格式为dd/MM/yyyy HH:mm:ss,所属时区为Australia/Melbourne:
场景1:DataFrame/Dataset列转换(主流Spark开发场景)
直接使用Spark内置函数实现,性能最优,适合批量数据处理:
import org.apache.spark.sql.functions._ // 示例DataFrame,melbourne_time_str为存储原始时间字符串的列 val df = Seq("21/10/2021 15:15:28").toDF("melbourne_time_str") val resultDf = df // 第一步:将字符串格式转换为Spark timestamp类型,指定原始时间匹配格式 .withColumn("melbourne_timestamp", to_timestamp(col("melbourne_time_str"), "dd/MM/yyyy HH:mm:ss")) // 第二步:将墨尔本时区的时间转换为UTC时间戳 .withColumn("utc_timestamp", to_utc_timestamp(col("melbourne_timestamp"), "Australia/Melbourne")) // 可选:如果需要输出固定格式的UTC字符串,新增格式化步骤 .withColumn("utc_time_str", date_format(col("utc_timestamp"), "yyyy-MM-dd HH:mm:ss"))
转换后示例结果:原始墨尔本时间21/10/2021 15:15:28对应UTC时间为2021-10-21 04:15:28,内置函数会自动处理夏令时偏移,无需手动计算。
场景2:RDD算子内单条数据处理
适合在map/flatMap等算子中处理单个时间字符串,使用Java 8+线程安全的时间API实现:
import java.time.{LocalDateTime, ZoneId, ZoneOffset} import java.time.format.DateTimeFormatter // 全局定义格式化器和时区,避免算子内重复创建 val inputFormatter = DateTimeFormatter.ofPattern("dd/MM/yyyy HH:mm:ss") val melbourneZone = ZoneId.of("Australia/Melbourne") val outputFormatter = DateTimeFormatter.ofPattern("yyyy-MM-dd HH:mm:ss") // 单条时间转换逻辑 val inputStr = "21/10/2021 15:15:28" val utcTimeStr = LocalDateTime.parse(inputStr, inputFormatter) .atZone(melbourneZone) .withZoneSameInstant(ZoneOffset.UTC) .format(outputFormatter)
注意事项
- 禁止使用线程不安全的
SimpleDateFormat类,分布式场景下会出现概率性的时间计算错误 - 时区必须使用标准区域ID
Australia/Melbourne,不要用AEST之类的时区缩写,缩写无法自动适配夏令时切换 - 上述Spark内置函数兼容Spark 2.3+及3.x全版本,低版本Spark可自行封装UDF用Java时间API实现转换逻辑
内容的提问来源于stack exchange,提问作者Arjunlal M.A
相关产品推荐
相关产品推荐

