Spark 2.3下实现Timestamp减去1毫秒的替代方案求助
Got it, let's work through this since Spark 2.3 doesn't support those handy sub-second interval operations that landed in 2.4+. Your source data is a string in yyyy-MM-dd'T'HH:mm:ssZ format, and you need to subtract 1ms then output it back as a string. Here are a couple of solid, tested approaches:
Approach 1: Use Built-in Spark Functions (No UDF Needed)
We can leverage unix_timestamp to convert the string to seconds, scale it to milliseconds, adjust by -1, then convert back to the required string format. This avoids custom code and works within Spark's native functions:
import org.apache.spark.sql.functions._ val adjustedDf = df.withColumn("adjustedTimestamp", concat( // Convert string to seconds, scale to ms, subtract 1, then convert back to formatted string with ms from_unixtime( (unix_timestamp(regexp_replace(col("sourceStartTimestamp"), "Z$", ""), "yyyy-MM-dd'T'HH:mm:ss") * 1000 - 1) / 1000.0, "yyyy-MM-dd'T'HH:mm:ss.SSS" ), // Add back the UTC 'Z' suffix lit("Z") ) )
How this works:
regexp_replacestrips the trailingZfrom the input string (sinceunix_timestampdoesn't handle that suffix directly).unix_timestampconverts the cleaned string to a Unix timestamp in seconds.- Multiply by 1000 to get milliseconds, subtract 1, then divide by 1000.0 to keep the fractional second value.
from_unixtimeformats the adjusted timestamp to include milliseconds (SSSpattern).concatadds theZsuffix back to match your original format.
Approach 2: Custom UDF with Java Date/Time APIs
If you prefer more control (or need to handle edge cases with more flexibility), a UDF using Java's date parsing/formatting works reliably. This is especially useful if your input might have edge cases like midnight rollovers:
import java.text.SimpleDateFormat import java.util.{Date, TimeZone} import org.apache.spark.sql.functions.udf // Define a UDF that takes the input timestamp string and returns the adjusted string val subtractOneMs = udf((inputTs: String) => { // Input formatter matches your source format (handles the 'Z' suffix) val inputFormat = new SimpleDateFormat("yyyy-MM-dd'T'HH:mm:ss'Z'") inputFormat.setTimeZone(TimeZone.getTimeZone("UTC")) // Critical: ensure we parse in UTC // Parse the string to a Date object, subtract 1ms val originalDate = inputFormat.parse(inputTs) originalDate.setTime(originalDate.getTime() - 1) // Output formatter includes milliseconds and adds back the 'Z' val outputFormat = new SimpleDateFormat("yyyy-MM-dd'T'HH:mm:ss.SSS'Z'") outputFormat.setTimeZone(TimeZone.getTimeZone("UTC")) outputFormat.format(originalDate) }) // Apply the UDF to your DataFrame val adjustedDf = df.withColumn("adjustedTimestamp", subtractOneMs(col("sourceStartTimestamp")))
Key Notes:
- Always set the timezone to UTC when parsing/formatting, since your input uses
Z(which denotes UTC time). This prevents unexpected shifts due to Spark's default timezone. - Both approaches handle edge cases like
2021-07-09T00:00:00Z→2021-07-08T23:59:59.999Zcorrectly. - The UDF approach is more readable if you need to extend this logic later (e.g., adjust by different intervals).
内容的提问来源于stack exchange,提问作者SteveTR

