如何将Spark DataFrame的秒级整数time列转换为HH:MM:SS格式?
Got it, so you have a Spark DataFrame where the time column stores total seconds as an integer, and you want to turn that into a clean HH:MM:SS string. Here are two straightforward, efficient approaches using Spark's built-in functions:
Approach 1: Use Date/Time Functions (Concise & Optimized)
Spark's from_unixtime converts epoch seconds to a timestamp, and date_format lets you extract just the time portion in your desired format. Even though your time value isn't epoch time, treating it as such works perfectly for calculating elapsed time:
import org.apache.spark.sql.functions.{date_format, from_unixtime} // Assuming your original DataFrame is named 'df' val dfWithFormattedTime = df.withColumn( "time_formatted", date_format(from_unixtime($"time"), "HH:mm:ss") )
For your sample value 29220, this will output 08:07:00—exactly 8 hours, 7 minutes, and 0 seconds.
Approach 2: Manual Calculation (For Custom Control)
If you want to break down the time components explicitly (maybe you need hours/minutes/seconds as separate columns later), you can compute each part individually and pad with leading zeros to ensure two-digit formatting:
import org.apache.spark.sql.functions.{concat, lpad, floor, mod, lit} val dfWithFormattedTime = df .withColumn("hours", floor($"time" / 3600)) // Total hours .withColumn("minutes", floor(mod($"time", 3600) / 60)) // Remaining minutes .withColumn("seconds", mod($"time", 60)) // Remaining seconds .withColumn( "time_formatted", concat( lpad($"hours", 2, "0"), // Pad hours to 2 digits lit(":"), lpad($"minutes", 2, "0"), // Pad minutes to 2 digits lit(":"), lpad($"seconds", 2, "0") // Pad seconds to 2 digits ) ) .drop("hours", "minutes", "seconds") // Optional: Remove intermediate columns
Sample Output
Either method will transform your sample data into this:
+-----+------+--------+-----+--------+--------+--------------+ | time|nodeID| posX | posY| speed | date | time_formatted| +-----+------+--------+-----+--------+--------+--------------+ |29220| 148 |5316.2 |500.0|59.04 |7/5/2010|08:07:00 | +-----+------+--------+-----+--------+--------+--------------+
Quick Notes
- Don't forget to import the required Spark functions at the top of your code.
- Approach 1 is usually better for performance since it uses Spark's optimized date handling.
- If your
timevalues go over 86400 seconds (one full day),HHwill show values like25for 25 hours. If you want to wrap around to 01 instead, usekk:mm:ss—but that's rare for elapsed time use cases.
内容的提问来源于stack exchange,提问作者inama dev

