如何将PySpark DataFrame中的时间戳列截断到日期级别?
Hey there! Let's figure out why you're getting null values instead of the expected truncated dates when using trunc(dt, 'day') in PySpark.
The Root Cause
The trunc() function in PySpark only works as intended with TimestampType or DateType columns. If your dt column is stored as a string (even if it looks like a valid timestamp), trunc() won't recognize the format and will return null—that's exactly what's happening here.
Step-by-Step Solution
Let's fix this by ensuring your column is properly typed first, then truncating to the start of the day:
Check your column's data type
Run this to confirm ifdtis a string:df.printSchema()You’ll likely see
dt: string (nullable = true)in the output.Convert the string column to TimestampType
Useto_timestamp()with your timestamp format (yyyy-MM-dd HH:mm:ss) to parse the string into a proper timestamp:from pyspark.sql import functions as F # Add a new timestamp column (or overwrite the existing one if you prefer) df = df.withColumn("dt_timestamp", F.to_timestamp("dt", "yyyy-MM-dd HH:mm:ss"))Truncate to the start of the day
Now you can usetrunc()(ordate_trunc()for more explicit syntax) on the timestamp column:# Using trunc() df = df.withColumn("day", F.trunc("dt_timestamp", "day")) # OR using date_trunc() (same result, more descriptive) # df = df.withColumn("day", F.date_trunc("day", "dt_timestamp"))
One-Liner Alternative
If you want to skip the intermediate column and do it all in one step:
df = df.withColumn("day", F.trunc(F.to_timestamp("dt", "yyyy-MM-dd HH:mm:ss"), "day"))
Verify the Result
After running the above code, your day column will show the expected values: 2018-04-07 00:00:00 and 2018-03-06 00:00:00.
内容的提问来源于stack exchange,提问作者Jared

