如何在PySpark中创建不含秒数的Timestamp类型列?
在PySpark中创建不含秒数的Timestamp类型列
问题场景
现有字符串类型的时间列,格式为yyyy-MM-dd HH:mm,希望将其转换为Timestamp类型,同时在显示时去掉秒数部分,最终输出如下:
+---+----------------+ | id| ts| +---+----------------+ | 1|2023-01-22 09:00| | 2|2023-09-11 00:09| +---+----------------+
且列的Schema保持为:
root |-- id: integer (nullable = false) |-- ts: timestamp (nullable = true)
原代码存在的问题:执行withColumn(...).show()会将show()的返回值(None)赋值给main_df2,导致后续调用printSchema()报错;同时直接转成Timestamp后默认会显示秒数。
解决方案
要实现需求,需要分两步:
- 将字符串时间转为Timestamp类型,并截断到分钟级别(确保秒数为0)
- 在显示时指定时间格式,避免默认显示秒数
完整代码示例
from pyspark.sql import SparkSession from pyspark.sql.types import StructType, StructField, IntegerType, StringType from pyspark.sql.functions import col, to_timestamp, date_trunc, date_format # 初始化SparkSession spark = SparkSession.builder.appName("TimestampWithoutSeconds").getOrCreate() # 构造测试数据 data = [(1,'2023-01-22 09:00'),(2,'2023-09-11 00:09')] schema = StructType([ StructField("id", IntegerType(), False), StructField("ts", StringType(), True) ]) main_df = spark.createDataFrame(data, schema) # 1. 将字符串转为Timestamp并截断到分钟(秒数设为0) main_df = main_df.withColumn( 'ts', date_trunc('minute', to_timestamp(col('ts'), "yyyy-MM-dd HH:mm")) ) # 查看Schema,确认ts为Timestamp类型 main_df.printSchema() # 输出: # root # |-- id: integer (nullable = false) # |-- ts: timestamp (nullable = true) # 2. 显示时格式化输出,去掉秒数 main_df.select( col('id'), date_format(col('ts'), "yyyy-MM-dd HH:mm").alias('ts') ).show(truncate=False) # 输出: # +---+----------------+ # |id |ts | # +---+----------------+ # |1 |2023-01-22 09:00| # |2 |2023-09-11 00:09| # +---+----------------+
关键说明
date_trunc('minute', ...):将时间截断到分钟级别,内部存储的Timestamp会把秒和毫秒设为0,保证数据的准确性。date_format(col('ts'), "yyyy-MM-dd HH:mm"):仅在显示时将Timestamp格式化为不含秒的字符串,列的实际存储类型仍为Timestamp,不影响后续的时间计算操作。- 注意不要将
show()的结果赋值给DataFrame变量,show()仅用于打印输出,无返回值。
内容的提问来源于stack exchange,提问作者bigdataadd
相关产品推荐
相关产品推荐

