PySpark中如何将毫秒级时间戳转为yyyy-MM-dd HH:mm:ss.SSS格式
解决PySpark中毫秒级Unix时间戳转yyyy-MM-dd HH:mm:ss.SSS格式的问题
你的问题核心是from_unixtime函数仅接受秒级Unix时间戳作为输入,而你传入了毫秒级数值,相当于把时间刻度放大了1000倍,才会出现几万年后的错误日期。以下是两种可行的解决方法:
方法一:直接对毫秒时间戳做除法后使用from_unixtime
将毫秒级时间戳除以1000转换为带小数的秒级时间戳,再传入from_unixtime并指定包含毫秒的格式:
from pyspark.sql.functions import col, from_unixtime srcdf = srcdf.withColumn("createdTime", from_unixtime(col("createdTime") / 1000, "yyyy-MM-dd HH:mm:ss.SSS"))
对示例值1696012940095来说,除以1000后得到1696012940.095,转换后即可得到正确结果2023-09-29 18:42:20.095。
方法二:先转Timestamp类型再格式化
先使用to_timestamp将转换后的秒级时间戳转为Spark的Timestamp类型,再用date_format指定输出格式,这种方式更贴合Spark的类型体系:
from pyspark.sql.functions import col, to_timestamp, date_format srcdf = srcdf.withColumn("createdTime", date_format(to_timestamp(col("createdTime") / 1000), "yyyy-MM-dd HH:mm:ss.SSS"))
额外注意点
如果你的createdTime字段是字符串类型,需要先将其转为数值类型再做除法:
srcdf = srcdf.withColumn("createdTime", from_unixtime(col("createdTime").cast("long") / 1000, "yyyy-MM-dd HH:mm:ss.SSS"))
内容的提问来源于stack exchange,提问作者Jatin
相关产品推荐
相关产品推荐

